Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:34.080509Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 33 inbound Pith citation observations for arXiv:2504.16074.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:34.080509Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:26.588538Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
37 of 37 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9fb619dc-0655-46f9-a089-71b7261d9dad · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Matharena: Evaluating llms on uncontaminated math competitions, February 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47808e1e-63b6-4379-be89-219c77e5a3fc · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Barnard, Gwen Clarke, and Nicholas Duncan
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 21281d2f-8d02-4c67-a64e-2376ac709fb2 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Claude 3.7 sonnet and claude code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8c60afe-df81-4e24-b795-e6b6b95374a9 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bad582-b991-49d5-a220-f9d2538e5fac · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Deepseek-v3 technical report, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b2ad016-0b4b-4d41-8ee9-4d8d2e563006 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e915b1-3f69-4335-bdaf-fd813bbf52cb · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0077a1a9-9f11-4e4d-b7f3-6e6f60b93cca · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Humanity’s Last Exam
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 08117895-fdc7-4b4d-a26d-cb6f88e7a82d · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Grok 3 beta — the age of reasoning agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c4fd2b8e-ca4a-49d0-8cf3-672212830c0c · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc1b6f5d-497f-4d1a-b22c-a15c876a78f4 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Aime 2024 dataset
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bde1cbdd-2d79-44d3-bdee-d98d79e6c54b · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Latex2sympy_extended package
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee9eee1d-086d-409b-83b4-a7115e296901 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Let’s verify step by step
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 582b2d8e-1e2b-4d65-b15a-67bae3492f85 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Smith, Mateusz Paprocki, Ondˇrej ˇCertík, Sergey B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ad35d8-1692-485e-b64e-b864fcca5b48 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models s1: Simple test-time scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e42cea65-ec8f-4b8d-a1b3-13a410475ee0 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9350666-a982-4fee-a586-f1346a2c0407 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models OpenAI o1 System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d048d44e-f5aa-442a-87df-c4efa157a6aa · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Learning to reason with llms, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f08bc57c-f69a-40ad-9320-397ddd8be7e6 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Introducing gpt-4.1
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65561f45-d099-440a-a4e3-ac47c85bd866 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Introducing openai o3 and o4-mini
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 58c4488e-6e5a-48a2-8ca1-587d0053cc8a · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Openai o3-mini: Pushing the frontier of cost-effective reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a67ee4c9-853c-4923-a60b-c51b62a5eb0b · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ed730f-4549-49b5-b9cd-f70eca1545c8 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cbd67fa0-d51c-4db7-ab79-e33e76f60d52 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62941eaf-0cf6-4105-b38f-956824500dc3 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503565a7-e63d-4c00-8349-55f9115d6ada · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwen2.5 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f83c6a-29e9-4236-b747-ac79206ef854 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabc71b7-3d63-4387-9711-29afd82208ef · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b239952-3c68-43ce-992c-2eb5c3372692 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6fa308a-aa40-4ab8-aa17-eed55df7368c · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models depolarization probability
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a72b832c-4db5-48b9-b544-6ac7cbe85804 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5075a008-77af-4fe4-b864-eb3d09c4eead · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f9779ada-3b7e-492f-a006-c64a9c51458c · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Use Ta = 300 K,ρs = 1000 kg m−3, ρa = 1.30 kg m−3,t = 100 nm, andg = 9.80 m s−2
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4223bd0-037f-446d-915b-0229e658390c · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models clustered mistakes
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54e5d61d-eee6-4113-bf31-bcf1d2a7dd26 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b3c9969e-bb01-4363-8deb-18c053c163a0 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b34ef4c4-0e83-4ccf-a245-aadfce01c791 · outbound
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models continue
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 34ceef5a-16a8-4f97-8686-d72cca1ab5ab · inbound
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · inbound
lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59b5a75-24ae-458c-8033-15a5d622bfe8 · inbound
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e72e41a-3861-42b4-b98c-1fab90fbf033 · inbound
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 137ad425-18a0-4f6a-bb26-72eb091899c9 · inbound
On Path to Multimodal Historical Reasoning: HistBench and HistAgent PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · inbound
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb23a208-8f20-4c44-afa0-73abc1de32fb · inbound
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef9be859-608e-4677-a58a-2d2f2b39bf22 · inbound
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ec1267-8c38-4f5b-8ca6-c35061de6a75 · inbound
SciDA: Scientific Dynamic Assessor of LLMs PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc12a3de-5e3b-4c8c-9447-ab0a6ee1263d · inbound
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 1949
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23f2682-f21a-46f2-be0b-74d2cd113e36 · inbound
Agentic Exploration of Physics Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b744c6-cc25-48c7-86ca-448723632ff8 · inbound
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cbac3c2e-1c69-46f2-8286-5d56524584fb · inbound
LLaDA2.0: Scaling Up Diffusion Language Models to 100B PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d603fdcf-2b82-42da-b44d-d8cc4cf1ac57 · inbound
Ministral 3 PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b7c46987-ace9-4a1a-9360-c9abda75a829 · inbound
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5682c563-bb91-4c9e-bcff-234fb65fec30 · inbound
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d0573e-e1ac-475f-a076-319faa5b08f4 · inbound
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation af40aa66-2d6c-4ea5-aac5-f4df7fab6785 · inbound
Vision Language Models Cannot Reason About Physical Transformation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 969d8e6d-5cbb-49a2-a9a8-34101c6babf9 · inbound
Seed1.8 Model Card: Towards Generalized Real-World Agency PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c20b9fcf-9eaf-4c23-ac30-c378aaf4a9db · inbound
PolyReal: A Benchmark for Real-World Polymer Science Workflows PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aaaff15b-3695-4c8d-a275-dc55b59c8288 · inbound
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f727223d-5410-41c5-ab9f-d9748d93aea9 · inbound
PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79a6c43f-2664-4ec9-b725-153a6c4c7309 · inbound
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5d636051-6d1a-4be7-abd3-9f50fa66a361 · inbound
Heterogeneous Scientific Foundation Model Collaboration PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c55c8f39-8847-494f-ac22-95dec3f160dd · inbound
PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 668e0963-bcde-40db-8433-fc7ccdcea260 · inbound
MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22bbed5b-1b7d-4d4c-87bf-c95da090f955 · inbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 188ce006-ec3f-4cf6-8db8-cef93e7c21fe · inbound
Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 323ff441-fde6-4d5e-95e1-41ece48248e6 · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2bb6da02-4f4e-4cd2-bfe9-f1a20a654948 · inbound
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1d9d6a85-fc24-471f-84df-0d55338e06e6 · inbound
Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f75970-1fd3-4dcf-9cb4-c4e41ddf764a · inbound
CLVisc Agent for autonomous relativistic hydrodynamics studies PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · inbound
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.