Pith. sign in

Paper Citation Record · LEDGER

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

As of 16 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 33 inbound Pith citation observations for arXiv:2504.16074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16074 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:34.080509Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:26.588538Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9fb619dc-0655-46f9-a089-71b7261d9dad · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.940582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.940582Z digest=sha256:80cc441a6278a3cc07d00b18caa15c2dd0e9cdbce95d46f914fa3cac8ed0f07b

Observation 47808e1e-63b6-4379-be89-219c77e5a3fc · outbound

This paper cites Barnard, Gwen Clarke, and Nicholas Duncan.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Barnard, Gwen Clarke, and Nicholas Duncan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.474908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.944742Z digest=sha256:d4f643ab737b633850509fe78a1e914391b4e1eae6b759d1e3f2e8017c9fc3ba

Observation 21281d2f-8d02-4c67-a64e-2376ac709fb2 · outbound

This paper cites Claude 3.7 sonnet and claude code.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Claude 3.7 sonnet and claude code

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.463383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.948615Z digest=sha256:5a757732b35e104701e76bfe7940a8e6912bf913d9e37858bc110bd8763b5f0b

Observation c8c60afe-df81-4e24-b795-e6b6b95374a9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.952739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.952739Z digest=sha256:e9def05d199dc32f2e48a65997c926ceffd4b15882375fc269e4c67f7026929e

Observation f5bad582-b991-49d5-a220-f9d2538e5fac · outbound

This paper cites Deepseek-v3 technical report, 2024.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Deepseek-v3 technical report, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.452093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.956459Z digest=sha256:e1b57a8f67a0f523b32fbf2455b2088031f7377c93272c96660278af272db3e5

Observation 9b2ad016-0b4b-4d41-8ee9-4d8d2e563006 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.960265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.960265Z digest=sha256:5a2656ebe5be04a0ed023e81d79a0e927f97f9d446fe71581c121ebe036697f6

Observation 84e915b1-3f69-4335-bdaf-fd813bbf52cb · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.964460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.964460Z digest=sha256:75e97207c262fd5742fc2b8a8aac5289fa9c9d72fa26eae7ffd554e8a0d0a8b7

Observation 0077a1a9-9f11-4e4d-b7f3-6e6f60b93cca · outbound

This paper cites Humanity’s Last Exam.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Humanity’s Last Exam

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.441328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.968871Z digest=sha256:ce6c2b4325191745eaa69f89661f46accdc46f7645cf7916b9c509f2fb90bc4b

Observation 08117895-fdc7-4b4d-a26d-cb6f88e7a82d · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Grok 3 beta — the age of reasoning agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.430619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.972315Z digest=sha256:b648b3aac16338b0e6549f6235d940d2635d53a4df0fd81e4b831ee27629f5d3

Observation c4fd2b8e-ca4a-49d0-8cf3-672212830c0c · outbound

This paper cites OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.975977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.975977Z digest=sha256:c54cd15145bf5268524b9ace40675bf8f3e3ef7c94e8c032b8d616fc012493ba

Observation fc1b6f5d-497f-4d1a-b22c-a15c876a78f4 · outbound

This paper cites Aime 2024 dataset.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Aime 2024 dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.412385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.979572Z digest=sha256:1f5ef62201b04f2e05846fcafa40f4457c90b64faae0b7efd4647d481200d611

Observation bde1cbdd-2d79-44d3-bdee-d98d79e6c54b · outbound

This paper cites Latex2sympy_extended package.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Latex2sympy_extended package

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.400976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:33.983147Z digest=sha256:ff8d8cdfee612587a1179716f421847c190b04e5220ea722199e1c8dd6cb156e

Observation ee9eee1d-086d-409b-83b4-a7115e296901 · outbound

This paper cites Let’s verify step by step.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Let’s verify step by step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.986589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.986589Z digest=sha256:eaa30aaf1685aad38feda4971afb4f54264b7a5f3aacad610e8af55fc3860fe8

Observation 582b2d8e-1e2b-4d65-b15a-67bae3492f85 · outbound

This paper cites Smith, Mateusz Paprocki, Ondˇrej ˇCertík, Sergey B.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Smith, Mateusz Paprocki, Ondˇrej ˇCertík, Sergey B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.989914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.989914Z digest=sha256:d63c0ca9026fabb616d0845d849df764c49757048c5c7b3739616fce973a7516

Observation 41ad35d8-1692-485e-b64e-b864fcca5b48 · outbound

This paper cites s1: Simple test-time scaling.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.993288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.993288Z digest=sha256:65fb7d9f8217f72e0006268758fce4ba177ad84d7e1a2d6abbbf1404a58db137

Observation e42cea65-ec8f-4b8d-a1b3-13a410475ee0 · outbound

This paper cites GPT-4o System Card.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:33.996697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:33.996697Z digest=sha256:e9786ea7b14791e1d3c023ec32380c4b068385b3a2d948c10fac6d3923b48fab

Observation c9350666-a982-4fee-a586-f1346a2c0407 · outbound

This paper cites OpenAI o1 System Card.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.000370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.000370Z digest=sha256:ee2e53f789a5e7a374b89c93b5ab958771bc41409de2b904660e0ca23c879e6a

Observation d048d44e-f5aa-442a-87df-c4efa157a6aa · outbound

This paper cites Learning to reason with llms, 2024.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Learning to reason with llms, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.375131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.004089Z digest=sha256:fe101ff196d3151e7eb841312aeaa6e7ebc958b33b9ce49fd473f7752bf63fa8

Observation f08bc57c-f69a-40ad-9320-397ddd8be7e6 · outbound

This paper cites Introducing gpt-4.1.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Introducing gpt-4.1

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.363151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.007688Z digest=sha256:07f5aa3f3136e90e05c51f90de7f8666abb2e2e78ce0a2c586218ebf8cb77b9c

Observation 65561f45-d099-440a-a4e3-ac47c85bd866 · outbound

This paper cites Introducing openai o3 and o4-mini.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Introducing openai o3 and o4-mini

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.352306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.011690Z digest=sha256:2723f07c88d419c6bef7a21443e2caf308a09522d43e43bd9cde6cdf27488d31

Observation 58c4488e-6e5a-48a2-8ca1-587d0053cc8a · outbound

This paper cites Openai o3-mini: Pushing the frontier of cost-effective reasoning.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Openai o3-mini: Pushing the frontier of cost-effective reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.341712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.015324Z digest=sha256:102d902f7e8961f113802b6d0471683a178a06d0759a63489a3ee250cdf89518

Observation a67ee4c9-853c-4923-a60b-c51b62a5eb0b · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.018693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.018693Z digest=sha256:6502b003ec55abe1771ae5da4a010b4ec5a42c5e099d734aef7c75fbb4eeb309

Observation 20ed730f-4549-49b5-b9cd-f70eca1545c8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.331030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.022284Z digest=sha256:ddbd14a6307df18fb294ede29ef1057b12cc8ce502f283420e322bf286862fcd

Observation cbd67fa0-d51c-4db7-ab79-e33e76f60d52 · outbound

This paper cites an unresolved cited work.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.025745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.025745Z digest=sha256:8951ddc50d2f951f75be0be3668274de78a606053cdac1d0dee8c3811239bf81

Observation 62941eaf-0cf6-4105-b38f-956824500dc3 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.029227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.029227Z digest=sha256:5d76e1c82da40d377ae3d3c6d1c2da775ebe582551228044975724d4e2829536

Observation 503565a7-e63d-4c00-8349-55f9115d6ada · outbound

This paper cites Qwen2.5 Technical Report.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.032695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.032695Z digest=sha256:103187bae7b27479bbf0883dfb03f1483cc9a5e06db054d3aaa02847402d15d2

Observation 99f83c6a-29e9-4236-b747-ac79206ef854 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2025.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.036047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.036047Z digest=sha256:252de80c503f7af6aade7e40e97d792ce46be2863d96362eefec2a50bdc92a4f

Observation eabc71b7-3d63-4387-9711-29afd82208ef · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.039955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.039955Z digest=sha256:994b8344856e8c397ac8199bd0f3e08f23d15151e841e25f22bd22249d84b839

Observation 9b239952-3c68-43ce-992c-2eb5c3372692 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:34.044550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:34.044550Z digest=sha256:419bf367cbf848d0a4411c54862b83fb445b94e0df59ed0b1a8107d09f18e9dc

Observation c6fa308a-aa40-4ab8-aa17-eed55df7368c · outbound

This paper cites depolarization probability.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models depolarization probability

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.300503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.049194Z digest=sha256:47f06deada55b58aff588da51d91120d608ac6c56558c590915396013efa26a7

Observation a72b832c-4db5-48b9-b544-6ac7cbe85804 · outbound

This paper cites an unresolved cited work.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:34.290184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.053956Z digest=sha256:ff92dade1e96194aebf7fc4c213429c92d7e961b897a5bcf74da75f8bdc8340b

Observation 5075a008-77af-4fe4-b864-eb3d09c4eead · outbound

This paper cites an unresolved cited work.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:34.278626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.058173Z digest=sha256:60dad39bc3a10ea0bfac18f6063c99a072663c1391408fd5d5350d3662ba70e7

Observation f9779ada-3b7e-492f-a006-c64a9c51458c · outbound

This paper cites Use Ta = 300 K,ρs = 1000 kg m−3, ρa = 1.30 kg m−3,t = 100 nm, andg = 9.80 m s−2.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Use Ta = 300 K,ρs = 1000 kg m−3, ρa = 1.30 kg m−3,t = 100 nm, andg = 9.80 m s−2

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.266168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.062616Z digest=sha256:5efbf9616ac3f282e687898eb0323c5326fca1af8a51d54d8ca1c53c5dc4da99

Observation e4223bd0-037f-446d-915b-0229e658390c · outbound

This paper cites clustered mistakes.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models clustered mistakes

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T11:16:34.255535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.067139Z digest=sha256:57740b2fe1406e9df281bab91264914c23db8567e83d876336df1e3a4c86fa93

Observation 54e5d61d-eee6-4113-bf31-bcf1d2a7dd26 · outbound

This paper cites an unresolved cited work.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:34.243637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.071795Z digest=sha256:889999b839bd5bb67899b6b1451d2d326ec692c378b0d5210a972f7d9ec3ad91

Observation b3c9969e-bb01-4363-8deb-18c053c163a0 · outbound

This paper cites an unresolved cited work.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:34.232507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.076402Z digest=sha256:7dd5e43d71081634dd679cf42c71401f4aad90ac9b60b73004e34d42d0be57e9

Observation b34ef4c4-0e83-4ccf-a245-aadfce01c791 · outbound

This paper cites continue.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models continue

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:34.221189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:16:34.080509Z digest=sha256:dcc187266e345f75255e08685ac66ee0ea9ec2319b4a87e09116b7e295a1e93d

Pith citing papers

Observation 34ceef5a-16a8-4f97-8686-d72cca1ab5ab · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:26.588538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:26.588538Z digest=sha256:f5f80d86792d0394f1d71134b0ac2813344150be21f439280534553323ac6d62

Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.524310Z digest=sha256:dd0cccdde7b4b911915582caa4fc259b12ee7fc1b5fe5153f5359e8e5ef48c8f

Observation f59b5a75-24ae-458c-8033-15a5d622bfe8 · inbound

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions cites this paper.

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:26.948825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:26.948825Z digest=sha256:b9280463f37e3227d377d1c45e731eb33b1dba3cf9eb237b587ecf0c8402acab

Observation 3e72e41a-3861-42b4-b98c-1fab90fbf033 · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:54.491719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:54.491719Z digest=sha256:a4229c3d86f7bbab98b67701986e1720f033c71441490a7363329f7faab019f6

Observation 137ad425-18a0-4f6a-bb26-72eb091899c9 · inbound

On Path to Multimodal Historical Reasoning: HistBench and HistAgent cites this paper.

On Path to Multimodal Historical Reasoning: HistBench and HistAgent PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:14.955222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:14.955222Z digest=sha256:eff39c5576e044be889864a960dcd17fdb85df3122407e5533575b8a957ad1bc

Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · inbound

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models cites this paper.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.013730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.013730Z digest=sha256:f3d4519709cbc423997922fc988b4b7b25373a98114105219eadff692829d629

Observation cb23a208-8f20-4c44-afa0-73abc1de32fb · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.713393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.713393Z digest=sha256:fa5bf7a8654156b94de73d7a3adaa0d3622669950ceb209e4b380ccbe668cfe9

Observation ef9be859-608e-4677-a58a-2d2f2b39bf22 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.353897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.353897Z digest=sha256:b6ccfc4b1116bcb2a2191b7f64ed8c73e8de2c9e1fe67941b5a15057947cc6e6

Observation a7ec1267-8c38-4f5b-8ca6-c35061de6a75 · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:18.603262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:18.603262Z digest=sha256:5227181d5f1bdf902fc10b2b1c2eb557cce7775243b6919dfeb2d1fccdb27f30

Observation bc12a3de-5e3b-4c8c-9447-ab0a6ee1263d · inbound

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems cites this paper.

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 1949

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:46.994892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:46.994892Z digest=sha256:924c6e030be3227ccd5dcf4f293edfc88af6ff3b5416a5dea67aefbfe3f13220

Observation f23f2682-f21a-46f2-be0b-74d2cd113e36 · inbound

Agentic Exploration of Physics Models cites this paper.

Agentic Exploration of Physics Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:31.512492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:31.512492Z digest=sha256:df5b77ab9853c1e98b60c3a1a84753fa02bfab40b739118ba10b8f48b9fc1d33

Observation e0b744c6-cc25-48c7-86ca-448723632ff8 · inbound

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark cites this paper.

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.306254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T11:52:10.205796Z digest=sha256:bfdbbf4cf2d79727fb7da5200cb1b1f8a36330e7ac88660b61ab053580a40843

Observation cbac3c2e-1c69-46f2-8286-5d56524584fb · inbound

LLaDA2.0: Scaling Up Diffusion Language Models to 100B cites this paper.

LLaDA2.0: Scaling Up Diffusion Language Models to 100B PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:53:21.116805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T18:53:20.911374Z digest=sha256:01d3fb260c9f738ab6372e0cf0db5fd417ff57a1f6a5db1a2ab641b83174cd97

Observation d603fdcf-2b82-42da-b44d-d8cc4cf1ac57 · inbound

Ministral 3 cites this paper.

Ministral 3 PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:24.750113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:e260daa7d788abe2004af52dc69d379db45fa9ff69890d1ae280ed6e6c874986

Observation b7c46987-ace9-4a1a-9360-c9abda75a829 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.220461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:24056c8b5cc43193b28ea979ee296cffec6de3b6c81e28d6856376f65154b5d9

Observation 5682c563-bb91-4c9e-bcff-234fb65fec30 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:24.771643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:24.771643Z digest=sha256:ee1aa496c4ed89e65db49bae10a76fdc244fbaf9ba165d5b33e77b4a10217c27

Observation d9d0573e-e1ac-475f-a076-319faa5b08f4 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:10:43.156734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:aeea6859f5ba434799cea463f012ab639a7314cf3da2b27df3088d7807f910ac

Observation af40aa66-2d6c-4ea5-aac5-f4df7fab6785 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:14d02bcefaaeb475a6d544e0c64c33c4da7d0c5dab603532b6e6a4322d023f3e

Observation 969d8e6d-5cbb-49a2-a9a8-34101c6babf9 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:45:14.405531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:863c8f896d42b00962d7a776773359244ebb83efa565825256a3712dae0927e7

Observation c20b9fcf-9eaf-4c23-ac30-c378aaf4a9db · inbound

PolyReal: A Benchmark for Real-World Polymer Science Workflows cites this paper.

PolyReal: A Benchmark for Real-World Polymer Science Workflows PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:58:12.284099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T19:55:17.366868Z digest=sha256:c760f08057a46a41959c625e3b15a588f725f451ad8b5bbb0602fffec4376fa4

Observation aaaff15b-3695-4c8d-a275-dc55b59c8288 · inbound

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research cites this paper.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.381486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:160d0439b0cfebc48f8f31829f74f31137e5d08099949687e3827bd07ff331ca

Observation f727223d-5410-41c5-ab9f-d9748d93aea9 · inbound

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement cites this paper.

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:13.110357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T06:13:18.967603Z digest=sha256:4e4aabaa254c902117734b1de7348dc773b7054ef69be6e09904ae873e3ac3fc

Observation 79a6c43f-2664-4ec9-b725-153a6c4c7309 · inbound

PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model cites this paper.

PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:26.628602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T03:43:18.640857Z digest=sha256:6e9558f95d9d7296d4e7517a9d185858d5526f808ea718ea5f85536e4fcad5a2

Observation 5d636051-6d1a-4be7-abd3-9f50fa66a361 · inbound

Heterogeneous Scientific Foundation Model Collaboration cites this paper.

Heterogeneous Scientific Foundation Model Collaboration PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:30:10.962510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T08:50:05.980191Z digest=sha256:8c2a3030c878d2cd2e8334adc585eec40925305cb0b27730cc0bd3c87d276e02

Observation c55c8f39-8847-494f-ac22-95dec3f160dd · inbound

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation cites this paper.

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:26.675661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T02:36:40.696567Z digest=sha256:1bab53c672244292025932e228e60354540df420a7af5d8c712e9a1ce003f395

Observation 668e0963-bcde-40db-8433-fc7ccdcea260 · inbound

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics cites this paper.

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:16:24.687383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T12:19:27.221596Z digest=sha256:dcde3a70346b0a8dbf5c4d1481ab909bae95d41a2000cab6973a2d1cdf89f4b9

Observation 22bbed5b-1b7d-4d4c-87bf-c95da090f955 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.619660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f921505041d09e63011330823e53f09b85da78b93ebe83ee72e7b91e510b5543

Observation 188ce006-ec3f-4cf6-8db8-cef93e7c21fe · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.878722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:78c3ffe698d6c68f19ddb449c9aa3e15aebb5e98c92e33e3e9156276f3aad020

Observation 323ff441-fde6-4d5e-95e1-41ece48248e6 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.387781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:cd443fcedbad0a461da1427a2d61c8c72304930433d651f5087b267d8f09c51f

Observation 2bb6da02-4f4e-4cd2-bfe9-f1a20a654948 · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.703824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:5c3ca13a416b9e23bd063fd64ae6ee66056d60ac6783398f8727d126a6f02836

Observation 1d9d6a85-fc24-471f-84df-0d55338e06e6 · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:fc9867b71940145fc9d96f3611be189e19d197a4f306522471ab0d6315d6d37c

Observation a9f75970-1fd3-4dcf-9cb4-c4e41ddf764a · inbound

CLVisc Agent for autonomous relativistic hydrodynamics studies cites this paper.

CLVisc Agent for autonomous relativistic hydrodynamics studies PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:52:32.240795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:52:32.240795Z digest=sha256:f5268b85e9a7838312d03c64dfb7d94c3c544a188bd61d7835f41663989920a1

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:63399343e146aaf8b1857fc3c8d1e5b8fc64c6fbaa29acb07941255c9a9ee0e4