Pith. sign in

Paper Citation Record · LEDGER

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2505.24823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24823 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:22:26.217604Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a48d0e-96b2-47ab-837b-9a9a873b4ed1 · outbound

This paper cites A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.263718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.263718Z digest=sha256:c131b3a644fa4311b74c2fa88c457e1b796384ed67bceea4574da7dc78b6c689

Observation 808a5c32-b05e-48b9-99ca-4c0c0033234a · outbound

This paper cites Romera-Paredes, M.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Romera-Paredes, M

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.403429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.403429Z digest=sha256:392ba41719ab2a9f51f406de9fa93901a4911ffde972eec08649a84867f68335

Observation fe59b0df-1003-4619-bb2e-5994d0365be5 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.541012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.541012Z digest=sha256:63c88f36f04cdec9e1adf8ccaee5e201cf810a60c20e736cd6b3f065a4f33531

Observation df50c6f1-8be3-4d9d-915d-7c18ab264605 · outbound

This paper cites Llm-sr: Scientific equation discovery via programming with large language models,.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Llm-sr: Scientific equation discovery via programming with large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.507274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:22.670967Z digest=sha256:ffce7041e5008eb2da7d2b97e0aa83056ea0816479c7ed65b7dcd0cfdf99b90f

Observation a364e9d4-33f0-4d5c-93a0-a456a9d46440 · outbound

This paper cites Brenner, and Eun-Ah Kim.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Brenner, and Eun-Ah Kim

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.854654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.854654Z digest=sha256:c9b54dafa913034c1b9011f8ce5c44db9b405713d48fb875932ff0e8fb5903ca

Observation 51298dbc-c492-4a51-8376-bb4d517b523d · outbound

This paper cites LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.956783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.956783Z digest=sha256:d5e0c44fc02417eacb8a911fbff6d93be699800af9c23139e5b68c4d129ff2c0

Observation 32b624ba-dac3-4396-a211-711b3a7bd2a1 · outbound

This paper cites Large physics models: Towards a collaborative approach with large language models and foundation models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Large physics models: Towards a collaborative approach with large language models and foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.042992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.042992Z digest=sha256:8661c71d842e62ffee6c50788598d071694f6d04829fc4123693975f8077f00a

Observation c4563436-9104-4f8f-87b4-a07a5a2783fb · outbound

This paper cites Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.142195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.142195Z digest=sha256:8a24e373f3405493b4a2f97ec008bd0a9fcd03a276c7e811f0ad46ded62173e6

Observation 0b2493fa-4321-4818-a7ef-0128d76ea2ca · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini Robotics: Bringing AI into the Physical World

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.253023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.253023Z digest=sha256:d43bc4b06a1217405bec602c3c9990584edfb34888db5e9c74bc38dec777368a

Observation 396fc029-6d63-4220-a6ad-614e4ce1b101 · outbound

This paper cites Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.366842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.366842Z digest=sha256:cac40650e3b027f0649c3cbbfd1da2b7e8d24392126784b2838cc95391dd4ab4

Observation 3f783fd1-951f-41cb-9688-782fa1291086 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.478351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.478351Z digest=sha256:f54567dd57de2cd4d8ae87419e7890e5eef765666cc477ed143504e0d094f825

Observation 803b530e-0457-400d-a200-728c2ce4442f · outbound

This paper cites Expert and novice performance in solving physics problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Expert and novice performance in solving physics problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.365459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:23.556910Z digest=sha256:aa5aba1e7c018367c4944b5734d9f7d89190bbc5afc5420aec65fc10b672a55b

Observation 9371a1df-69bf-44c3-85d2-988571c33163 · outbound

This paper cites an unresolved cited work.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.661672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.661672Z digest=sha256:65f3fb66c7dcff1c01ab59039695b674f1823aa9e376bfce5d0456a2de255501

Observation f7c2d1ae-54f7-48bf-9368-62c51f1b0030 · outbound

This paper cites Cognitive load during problem solving: Effects on learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Cognitive load during problem solving: Effects on learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.740680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.740680Z digest=sha256:1d2fce325cd44f8ff41a3d0dad2712573316dcab57359f465727a51f697b0ee2

Observation 4b35877a-f8ac-416f-aebc-eba31f4ffe22 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.833827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.833827Z digest=sha256:8d7231af3d0ee601c74639a33c36900badcfa33bcb1a7382502d23e44ac64d65

Observation 521599b2-9065-4a7f-931a-2bbcfdfd0ee4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.924503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.924503Z digest=sha256:1253ba29f7c42b2e162e3309f5e9d1addfb8e594586504f1f15d772ee000e36a

Observation dbb57740-cea8-471f-bcbf-fc7a73709532 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.996077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.996077Z digest=sha256:4fd05a22c234b23ab3e64570a37a284a4a49c92d1f86524e1490fae7cf90bcc5

Observation f4e69425-a990-4618-bc0a-e1b6462b638a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.060193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.060193Z digest=sha256:5da21f89b2953093c3e4086cafaff620db327faef750f9f6c8bde6dd93bec5f2

Observation cdf12970-07b3-43f0-bc47-85ff7762fa00 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.129251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.129251Z digest=sha256:53026862c825b087df7fd7e83c5bbe5c97c3a51b49da6a37c2e337aa40925aed

Observation 79cc130c-bbb1-43d2-b840-5dcaa3d9d231 · outbound

This paper cites Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.207587Z digest=sha256:bee1633139812c15dc2e3aa689443e5ab08a4b3789fee32c71a414cc05a3a0d7

Observation 02dda92c-2780-45b9-8f2c-1f865cb1b5ae · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.297967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.297967Z digest=sha256:e356c0558a3aaef5dd908b6ff5ba2be58302b76a4ad3b29d6a028d8c779ce71d

Observation 0c387862-5d82-4e16-9e86-65313924be74 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.372229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.372229Z digest=sha256:707a80b5a0d8080494ebb5145ba65a77cc02722ada850e74729ce1bb3f1c61de

Observation a869ab84-d082-48b6-ba44-90cd9788c3f8 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Scieval: A multi-level large language model evaluation benchmark for scientific research

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.130295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:24.440236Z digest=sha256:9d30e13a6ab633ffca61f0ad9a965acf32e648bdcf864df8427ac44b552e6f59

Observation 5a431d61-daf8-4a7e-897a-e5c964430a73 · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models TheoremQA: A Theorem-driven Question Answering dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.562397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.562397Z digest=sha256:db9138c485bd28b52a41968f67a73e80182ae751fcd7f66ff1c16a090885f8cb

Observation af06dde0-a12c-42a3-ade6-77517c781a61 · outbound

This paper cites Humanity's Last Exam.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Humanity's Last Exam

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.630990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.630990Z digest=sha256:573e04553054c23b81b51d067273397877bf6f72b086705bbbf2e8dbfcecbc49

Observation 4a96a0a4-9d2a-495f-b9bc-2f049d547835 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.704748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.704748Z digest=sha256:77c65b49952140dd3663731bf97338912e2d3843e30d0f902fcfbf58065d135a

Observation 180c217d-8b14-420f-b85d-f8106afa1290 · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.776896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.776896Z digest=sha256:040df4100fd9dee5705d96186d68566c45b4adaa9247f932a17ae0d56be082b4

Observation 9b4783e0-6f33-4c67-a1ba-a05baec08497 · outbound

This paper cites Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.850334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.850334Z digest=sha256:925b3b61bf48dccb27283d2c6c606ba5f1bee8d034638e76b981f9d074b6d097

Observation 03c59c30-0f57-4c2b-9221-927d494fe54b · outbound

This paper cites UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.926698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.926698Z digest=sha256:f8a4c881c26a3563c3491f68c52ebfde3ef9ba112fcb93c215a4f91dde766936

Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.013730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.013730Z digest=sha256:f3d4519709cbc423997922fc988b4b7b25373a98114105219eadff692829d629

Observation aa447c21-1001-482c-9c3f-e2f2bc2942c1 · outbound

This paper cites PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.112064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.112064Z digest=sha256:b2fcce9df2bf3a6209eda7347bb7e03c54002ef7635163761754c3dc8b51dbe9

Observation 9993ea64-1444-49b2-b6e0-7d4512665de9 · outbound

This paper cites CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.206982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.206982Z digest=sha256:540d8c007eebdd1dc5b0b0151608ed164c01628064b00ca1d6fc069611f71467

Observation 18e58079-ba3c-4c38-ab28-ba13d822cc40 · outbound

This paper cites Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.273400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.273400Z digest=sha256:adc8e01f6ffc58011ab0a9e51e0b49bf790fbc489c90c7dfa00aa022c653657c

Observation 6da479d7-7839-4893-9e43-1d28264fe77c · outbound

This paper cites FEABench: Evaluating Language Models on Multiphysics Reasoning Ability.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.362149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.362149Z digest=sha256:8a237bcc86a8fa1c79c8987c221af0ee259dc14461ca1e6a74ef0422ef8c4765

Observation 7a961bc9-1382-43b1-9101-ff2fc5c1d320 · outbound

This paper cites Learning to reason with llms, September 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learning to reason with llms, September 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.886863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:25.433618Z digest=sha256:b4ea60dcdb8a35a5b244cee97fdb9a2ae11c47d44a0528a59aabd2bedf73bb5b

Observation f3766474-28fc-4d52-aaee-35044a14fcff · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.520633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.520633Z digest=sha256:58f5d3e709e04cc0c21cf92ac26ad2832624ff73baaee9f65db5921fcb478c8f

Observation 82d86910-6508-4c19-a1ed-f374995122c0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.610427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.610427Z digest=sha256:cc785d85634a21d3bde65eb935a2ad9dba6f8831baa05a42b672aff985301c85

Observation df088cbf-23f6-4b89-ad03-62d780ba7c78 · outbound

This paper cites Claude 3.7 sonnet and extended thinking mode.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and extended thinking mode

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.712882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:25.698133Z digest=sha256:248fd3b72bf53c8334b818aaafb26799f91aefcab6c09bc500b8e34ae6344dd4

Observation 9c98da01-6cb7-4932-8275-2d57c4e85b04 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.536957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:25.762655Z digest=sha256:00ba8a1bea8046fc476a705ea1d281190a6760ff2c9ec3049b37a19d2a570865

Observation 610037b8-ea61-4d33-9e03-36693ba95cde · outbound

This paper cites Openai o3 and o4-mini system card.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Openai o3 and o4-mini system card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.353683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:25.833242Z digest=sha256:2695f5920a94cef1ec72570e5e1422c378c2bb904aade91f462241060264d1fe

Observation bc3d9c65-2030-4145-a99d-14949198f77d · outbound

This paper cites URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:25.907140Z digest=sha256:4eda5f3f16be17d975d2b2723e244b1b76de69db9d2ab4660438e54955e35265

Observation 0c3101cb-d4ed-4803-9e0f-7821cf0c4c4f · outbound

This paper cites Introducing gpt-4.1 in the api, April 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Introducing gpt-4.1 in the api, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.980401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.980401Z digest=sha256:96bd052cd1a502c8df3ae4e303a859c4c91756926955824b828e872a8e8d366f

Observation 41e85d35-4d36-4afa-9448-d6f018c2cc8d · outbound

This paper cites Claude 3.7 sonnet and claude code, February 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and claude code, February 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.046469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:22:26.052171Z digest=sha256:6af6599ec2bbd8db028a931d65b4d088486dd465824e47a3470572563d4d5f71

Observation 2d41297a-7b22-48f6-824a-25225da34672 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:26.127078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.127078Z digest=sha256:1ee05e1dcf422fb73a5ef67bf85c0681c9f441f87fc9dbb1ba42ac80862ab9dc

Observation 56ffc97e-5c6e-4606-9adb-ba5a36adf4f6 · outbound

This paper cites DeepSeek-V3 Technical Report.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-V3 Technical Report

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:22:26.217604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.217604Z digest=sha256:52525ea84e2d762b3e206ab74086d04185ee236d3825b85088a668af6829a73b

Observation 0e7d5fea-6140-40d0-9f66-44aae2f38c47 · outbound

This paper cites LLM-SR: Scientific Equation Discovery via Programming with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-SR: Scientific Equation Discovery via Programming with Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.780339Z digest=sha256:1afbc05d3eb3295e3b9d1f234ab57f14926405c84f284cdf118d5fd791a7539f

Pith citing papers

No inbound Pith citation observations are available.