Pith. sign in

Paper Citation Record · LEDGER

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2505.24823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24823 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:22:26.217604Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a48d0e-96b2-47ab-837b-9a9a873b4ed1 · outbound

This paper cites A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.263718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.263718Z digest=sha256:6e8e8ef44a05e1592d29d991bf46e22147732093dfa7b52b1db7880882465fcc

Observation 808a5c32-b05e-48b9-99ca-4c0c0033234a · outbound

This paper cites Romera-Paredes, M.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Romera-Paredes, M

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.403429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.403429Z digest=sha256:317dd36e546f53e3043179d40ea2095c6114ed972b8059a4dba5ee25d188503a

Observation fe59b0df-1003-4619-bb2e-5994d0365be5 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.541012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.541012Z digest=sha256:f19a6e79f0e898e261a00b2794cd24252313bac8b4aa0a79259ae20329109873

Observation df50c6f1-8be3-4d9d-915d-7c18ab264605 · outbound

This paper cites Llm-sr: Scientific equation discovery via programming with large language models,.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Llm-sr: Scientific equation discovery via programming with large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.507274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:22.670967Z digest=sha256:2ef69c2619479e1124edca45864e79a6ee5c38ce4af45b8156a0da1fa42479f1

Observation a364e9d4-33f0-4d5c-93a0-a456a9d46440 · outbound

This paper cites Brenner, and Eun-Ah Kim.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Brenner, and Eun-Ah Kim

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.854654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.854654Z digest=sha256:9fb5657ac51a97fbe2acb7e28ba1d0ef0396887004ce0c3e13060f6aca03a775

Observation 51298dbc-c492-4a51-8376-bb4d517b523d · outbound

This paper cites LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.956783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.956783Z digest=sha256:5021d44781cf79dfc6f9ef18c9060d5381159365de5c68d3248f4156823baf38

Observation 32b624ba-dac3-4396-a211-711b3a7bd2a1 · outbound

This paper cites Large physics models: Towards a collaborative approach with large language models and foundation models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Large physics models: Towards a collaborative approach with large language models and foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.042992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.042992Z digest=sha256:6fee2f156c7399cf1b8d086fecc14d5fc35d7df74905719c2dd83b1e6cc2a5a3

Observation c4563436-9104-4f8f-87b4-a07a5a2783fb · outbound

This paper cites Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.142195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.142195Z digest=sha256:183915ddb9c7c6e2f7f79d7ed5dce4187d9ebc02af3d5615d3eb5877a957723d

Observation 0b2493fa-4321-4818-a7ef-0128d76ea2ca · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini Robotics: Bringing AI into the Physical World

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.253023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.253023Z digest=sha256:1ffce688f8f8043406f1a940e477145b325e0771fecab67303a20d50ab807d38

Observation 396fc029-6d63-4220-a6ad-614e4ce1b101 · outbound

This paper cites Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.366842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.366842Z digest=sha256:c984796326329b497cb2721876ac9b84791d54674fcd61b88e129b946f7df6b2

Observation 3f783fd1-951f-41cb-9688-782fa1291086 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.478351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.478351Z digest=sha256:796f56e03e5e1a622e6c7001bdf86794c72f9fb73f1e48142b8fa372e2144880

Observation 803b530e-0457-400d-a200-728c2ce4442f · outbound

This paper cites Expert and novice performance in solving physics problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Expert and novice performance in solving physics problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.365459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:23.556910Z digest=sha256:d9227a86e0b39e0eb3d3ba11fe0da5315cfb237e026e28743149646e795d7872

Observation 9371a1df-69bf-44c3-85d2-988571c33163 · outbound

This paper cites an unresolved cited work.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.661672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.661672Z digest=sha256:f39b2c72c731d38a2133b00fad7f7bfbd9889c8532ef417659297abec760808d

Observation f7c2d1ae-54f7-48bf-9368-62c51f1b0030 · outbound

This paper cites Cognitive load during problem solving: Effects on learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Cognitive load during problem solving: Effects on learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.740680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.740680Z digest=sha256:a08d685edd077620bc7441fb6eb44f7d4472d16bfef9c4729433939c696e26e3

Observation 4b35877a-f8ac-416f-aebc-eba31f4ffe22 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.833827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.833827Z digest=sha256:f75d4e8c1cd99120e30003963f291f01b542bc1ef450cc624a2caa287cdd4b53

Observation 521599b2-9065-4a7f-931a-2bbcfdfd0ee4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.924503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.924503Z digest=sha256:588f041d0a6c25f983959c0223fc172dce280aaf3410c596b8898ac2bf65405f

Observation dbb57740-cea8-471f-bcbf-fc7a73709532 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.996077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.996077Z digest=sha256:7f658421a326aafd65bf9b89a5fd082a6bdd7877a0e7862ec738fa12a2ac9122

Observation f4e69425-a990-4618-bc0a-e1b6462b638a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.060193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.060193Z digest=sha256:3806d2b333dc6c693a4d9a35f828f33f91a90e979e6b870fc137793521c5d883

Observation cdf12970-07b3-43f0-bc47-85ff7762fa00 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.129251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.129251Z digest=sha256:579ae6f37dc2c155c196d8025ef78957a49a626a981a3ae29c1bfaea8d24ee9c

Observation 79cc130c-bbb1-43d2-b840-5dcaa3d9d231 · outbound

This paper cites Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.207587Z digest=sha256:2ab79e2d73b1f906ff764944fa20e4898daf8442a8aef38cd77aa5543f7bfe9b

Observation 02dda92c-2780-45b9-8f2c-1f865cb1b5ae · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.297967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.297967Z digest=sha256:33c2a14f3a0b27fc143d555d5849e342a4c8ff3036ef75bf5c74d3f8e746b803

Observation 0c387862-5d82-4e16-9e86-65313924be74 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.372229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.372229Z digest=sha256:7edc5fac233096529248c1a30936b6be909c61a6bef037b4b947b2be20c9cf56

Observation a869ab84-d082-48b6-ba44-90cd9788c3f8 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Scieval: A multi-level large language model evaluation benchmark for scientific research

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.130295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:24.440236Z digest=sha256:ab9ddee25dd9a19ea4b1268cda8d001708225523b7823e0cffcd4a61bca2dd8a

Observation 5a431d61-daf8-4a7e-897a-e5c964430a73 · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models TheoremQA: A Theorem-driven Question Answering dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.562397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.562397Z digest=sha256:3ee1fe5550f3b62d3a963f4d3b51fe504589c58f2cf6ff7109a03cbf456cfe5b

Observation af06dde0-a12c-42a3-ade6-77517c781a61 · outbound

This paper cites Humanity's Last Exam.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Humanity's Last Exam

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.630990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.630990Z digest=sha256:5e8907fa82dca2fdd0d110ab9f6e19f80129d28bec29fb1e12847b0712e85522

Observation 4a96a0a4-9d2a-495f-b9bc-2f049d547835 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.704748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.704748Z digest=sha256:624fce67e718dda6c451295007eef9410ad762f1281532be9805d35bbf8d0f66

Observation 180c217d-8b14-420f-b85d-f8106afa1290 · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.776896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.776896Z digest=sha256:59960edc0f64ac35443011ab3ce55ab6f4a1e0826e2dcbf7750042cfd636f173

Observation 9b4783e0-6f33-4c67-a1ba-a05baec08497 · outbound

This paper cites Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.850334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.850334Z digest=sha256:d9fe400d26ce67a0943eac941ea09ff0ee711e53fff613f3a77a7a274fde4608

Observation 03c59c30-0f57-4c2b-9221-927d494fe54b · outbound

This paper cites UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.926698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.926698Z digest=sha256:9d3aec8e626aafc71a2a5a31dd7b09a791c353cbd8fe1f316533777779d41b71

Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.013730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.013730Z digest=sha256:e9194c940b665f85a2e0aec5975482fb0452cbfc9c2b1a813db075e8857839d3

Observation aa447c21-1001-482c-9c3f-e2f2bc2942c1 · outbound

This paper cites PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.112064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.112064Z digest=sha256:97dc4a6720c736adff2d04f3182c3845b5253a69f203f69c68e5e9a43ccd1b60

Observation 9993ea64-1444-49b2-b6e0-7d4512665de9 · outbound

This paper cites CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.206982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.206982Z digest=sha256:d307c0aa9ea274830ad0b4fb6e578fc0cc5031bc3210fbca3d6dbe802b9717bb

Observation 18e58079-ba3c-4c38-ab28-ba13d822cc40 · outbound

This paper cites Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.273400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.273400Z digest=sha256:bff25831fd129291b7b38ee796a8129eab7fb161b75159dd7f18228469a693a7

Observation 6da479d7-7839-4893-9e43-1d28264fe77c · outbound

This paper cites FEABench: Evaluating Language Models on Multiphysics Reasoning Ability.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.362149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.362149Z digest=sha256:d115f54e335f7535d50006d3365ec9da021a2ab4839e0f0806c7f409b6861c0e

Observation 7a961bc9-1382-43b1-9101-ff2fc5c1d320 · outbound

This paper cites Learning to reason with llms, September 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learning to reason with llms, September 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.886863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:25.433618Z digest=sha256:f462795428a58be26f71bdf12a6398f0cb3a9d086f42483a205e1032bf9454e6

Observation f3766474-28fc-4d52-aaee-35044a14fcff · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.520633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.520633Z digest=sha256:f7e30536f3ce00206efc0eb2f7486bb7ddee0fcab3479d0e6188c5f3aa438a2e

Observation 82d86910-6508-4c19-a1ed-f374995122c0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.610427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.610427Z digest=sha256:d0caeb5d5d8faa66a9a39102d3f44c75bdedae0e95d7244d8a980547aea950ab

Observation df088cbf-23f6-4b89-ad03-62d780ba7c78 · outbound

This paper cites Claude 3.7 sonnet and extended thinking mode.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and extended thinking mode

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.712882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:25.698133Z digest=sha256:310da93f7b0bc3eb1619c348bf2406398b9ec81b46b49c4135f51f7d55048f07

Observation 9c98da01-6cb7-4932-8275-2d57c4e85b04 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.536957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:25.762655Z digest=sha256:bbaadd72a8dc5624da1dd58fbf46874b29ca127f92f90670eb000d4abf94fc3b

Observation 610037b8-ea61-4d33-9e03-36693ba95cde · outbound

This paper cites Openai o3 and o4-mini system card.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Openai o3 and o4-mini system card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.353683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:25.833242Z digest=sha256:13e8dd1d900401c60f7d92e4a6e559bf4aa0f3f136b04502abc0041a934a1e65

Observation bc3d9c65-2030-4145-a99d-14949198f77d · outbound

This paper cites URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:25.907140Z digest=sha256:4da62fb75777b150f3da5f57968302eeaae049782c167b16ce2265ee590bd36c

Observation 0c3101cb-d4ed-4803-9e0f-7821cf0c4c4f · outbound

This paper cites Introducing gpt-4.1 in the api, April 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Introducing gpt-4.1 in the api, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.980401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.980401Z digest=sha256:7b92b8fe17bed1a15eeb59190b954860c00baed2c8d4f935bc0d7608e83832f9

Observation 41e85d35-4d36-4afa-9448-d6f018c2cc8d · outbound

This paper cites Claude 3.7 sonnet and claude code, February 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and claude code, February 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.046469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:22:26.052171Z digest=sha256:b996180c20a483c1f34947d7d01ca8f3eff67c7fb612247ba3d593ff0dd74a60

Observation 2d41297a-7b22-48f6-824a-25225da34672 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:26.127078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.127078Z digest=sha256:bd09256cdb177e1f2de9861b997f03167613e8492e6826893f20855801c98d44

Observation 56ffc97e-5c6e-4606-9adb-ba5a36adf4f6 · outbound

This paper cites DeepSeek-V3 Technical Report.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-V3 Technical Report

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:22:26.217604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.217604Z digest=sha256:7ea465c45ed541d2dd38ebd781209c39bc1341af5df4fed2c8cbfb88c7812cf8

Observation 0e7d5fea-6140-40d0-9f66-44aae2f38c47 · outbound

This paper cites LLM-SR: Scientific Equation Discovery via Programming with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-SR: Scientific Equation Discovery via Programming with Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.780339Z digest=sha256:ebe7e939cbd16f84ed34412942b45919a844fea6e512878cd563756d5ad53536

Pith citing papers

No inbound Pith citation observations are available.