Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards

As of 9 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2607.10474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10474 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T11:28:23.511747Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c2a8297-13ac-4bdb-9160-cda16609d97a · outbound

This paper cites American Mathematical Society, 2 edition, 2010.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards American Mathematical Society, 2 edition, 2010

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3b23ee564a8f791f056a32b11001e479b706f470fbdb51c0d6476ea85ecebd6c

Observation 3e264e67-a8dc-4868-a8b1-2be3021a48ab · outbound

This paper cites Cambridge university press, 2002.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Cambridge university press, 2002

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:950b3540e35789f07e91d696f358cd0c45475893f0c2399fb6b6e406fae2f860

Observation 36fc7b67-b486-4e7c-ac30-93b3ab78b628 · outbound

This paper cites Springer, 1994.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Springer, 1994

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:667a7b025ae75bb433115e379f24518114e5fb481757fe9c031b8ce932d72d78

Observation a5a69d94-9d79-4082-a741-b832f8b37626 · outbound

This paper cites SIAM, 2000.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards SIAM, 2000

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:648abff0773a40249cb5fd754cc135ac79080105380ba37ce6e5b568664b38f3

Observation 115fd3ce-6941-4823-a23e-6b213aeb88f2 · outbound

This paper cites SIAM, 1998.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards SIAM, 1998

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f3d32ef15786ef28b3a2246613e54797b1f6938ec634f43d6788ac15a287a3aa

Observation a1a414a4-0922-475a-ab4e-27c0163a690c · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3e0dffec5923756047964ffc96cae04833fd36b2ad588e0051d7492a0224faea

Observation 4b56acf4-4e06-4238-bcb4-8a5b77d36cda · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Measuring Coding Challenge Competence With APPS

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ec00e53df5df2f4f588a576f19343b513ebfbdb6ce343af13ef361e260f0b75c

Observation 37cda1a9-6b67-4a43-b293-550fd4b3bf9f · outbound

This paper cites Program Synthesis with Large Language Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:653d74f8480d32572bae212a0962f0ce1d69dad915d6c068453df79a949351ff

Observation cdccca16-61e0-4fb2-a2cc-cb9f75ad6180 · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:54a4cdecce0f963d7eadab077c793a5ece4064f856a635d13889878d2fcf8c54

Observation 130a9fb1-6740-4829-bbfe-fd750d70b16c · outbound

This paper cites Narasimhan, and Yuan Cao.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Narasimhan, and Yuan Cao

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:5dda6d5b2bc828c2038d8dc561029ee823dced9285dc4d7a4f90a536c9f6ced9

Observation 60040e24-02a3-4817-bac6-c7be406df154 · outbound

This paper cites URLhttps://openreview.net/forum?id=WE_vluYUL-X.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards URLhttps://openreview.net/forum?id=WE_vluYUL-X

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:0c9843ff2237ed04722185c5c1739660f4dad6f5ec2bfa3e6f90cf7a8034d728

Observation dfd47062-a949-4b64-969c-22bb38ce2cf8 · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:4a358e2b0d845ea7862596e613803baeb1ceaf0b08810b57dbe5b92882c39789

Observation af13fd51-764c-45c3-a680-4bcf05b56b9c · outbound

This paper cites Codepde: An inference framework for llm-driven PDE solver generation.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Codepde: An inference framework for llm-driven PDE solver generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:0f137d1f25eed39a2a3e71bf7f6800cee9d1be4a966401ea72283ff583ef9452

Observation 7606bdd5-7e76-4301-aae2-79bf3252bb7c · outbound

This paper cites SciML Agents: Write the Solver, Not the Solution.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards SciML Agents: Write the Solver, Not the Solution

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:03b4be12f91f5f1b205336cb79ee9f490b3f3f224a63d8285c9fa6b6cf82571a

Observation d109cd36-cea3-4614-a1b6-1b172d6f578c · outbound

This paper cites Autonumerics: An autonomous, pde-agnostic multi-agent pipeline for scientific computing.arXiv preprint arXiv:2602.17607, 2026.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Autonumerics: An autonomous, pde-agnostic multi-agent pipeline for scientific computing.arXiv preprint arXiv:2602.17607, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ff0de99d6ceb01060f4c224620dfa009e4c395022cdc7e167ed8d54eafc81341

Observation 89b9b175-75c9-4af5-8740-28784aeef19d · outbound

This paper cites All-fem: Agentic large language models fine-tuned for finite element methods.Computer Methods in Applied Mechanics and Engineering, 457:118985, 2026.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards All-fem: Agentic large language models fine-tuned for finite element methods.Computer Methods in Applied Mechanics and Engineering, 457:118985, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:7335554298a8f49b75e46823d8819297061daad9e81ff8138743b3a575d4ef15

Observation 01be39bb-1171-4877-ba82-54d12aebe090 · outbound

This paper cites Pde-agent: A toolchain-augmented multi-agent framework for pde solving.arXiv preprint arXiv:2512.16214, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Pde-agent: A toolchain-augmented multi-agent framework for pde solving.arXiv preprint arXiv:2512.16214, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:6f82407b894d8295707eac30bf9dd2d30df250f4dccdfcff74e47ae73f8fcbba

Observation fe4afabf-5ef0-477e-99b5-ba907efa9214 · outbound

This paper cites PINNsAgent: Automated PDE Surrogation with Large Language Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards PINNsAgent: Automated PDE Surrogation with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:53d57ce8e600c8274225ee481c98f695db741d958da2062f9d510c53c8ebe1e2

Observation 9df63f9d-7f74-4986-b992-8046fc8575f6 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:06da0ff16d92ec091e6e0e270e930a93341aae445264b953365edc72f81cd846

Observation 54306c0d-dc39-4df0-b4bd-4263ab5ae388 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:1f522d8ebbb91792c8d51db275e355c044fb001e13e7c79c6298a0be19feb2b6

Observation 624f6a03-86d6-4a14-a9db-71e2c052226b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:abe90d9a446deb256d33b448ad6d2978d8082ca59719e807c01f644045cc9a82

Observation 5bbf41bc-d83b-479f-bdd8-71f5097d3cdf · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f86ea8dcd46320be7a80b3cc789e8722865ecb5c62df40c37f8922c312ecf062

Observation 39fc4ebd-d77d-44b6-adb1-6c1a7d4acaf5 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Solving math word problems with process- and outcome-based feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:306b44e2c0c4a7d95b252d575282fa80be1eef50b107223e7064844932204bbd

Observation 4d811320-eb90-4cd9-a7f8-44200d21492a · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Efficient memory management for large language model serving with pagedattention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:be82f6024324e22625c7dae4ea8536c55240da821804e7f9fe4e03b6f00d1d7b

Observation e84096b0-83c2-4f0e-b77c-6b7afe6c34e1 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Hybridflow: A flexible and efficient rlhf framework

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:0aab157fcded7f6c98214120bfc43f188d449f80d9cfa6564bea9008ddbbd51f

Observation 45d4cca4-fa63-48b2-bc1a-8874f9d0d396 · outbound

This paper cites Huerta, and Hao Peng.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Huerta, and Hao Peng

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:7c52d649cb0f1ac26c5eefebff5809d2d116312c3936cad11b70e2e3c33fd0b9

Observation 0275feb3-b127-4e03-bd36-874c45a4be81 · outbound

This paper cites Foam-agent: Towards automated intelligent cfd workflows.arXiv preprint arXiv:2505.04997, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Foam-agent: Towards automated intelligent cfd workflows.arXiv preprint arXiv:2505.04997, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ae0c701f9455ceea80a568e5e07a24be9bddec831aa443a75e8001c410c3bedd

Observation a85e05d4-686d-4fb3-8def-9ed2414b218d · outbound

This paper cites Openfoamgpt: A retrieval-augmented large language model (llm) agent for openfoam-based computational fluid dynamics.Physics of Fluids, 37(3), 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Openfoamgpt: A retrieval-augmented large language model (llm) agent for openfoam-based computational fluid dynamics.Physics of Fluids, 37(3), 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:acf8ce3d427a6f9732aa0452b552e3088102e10402c223af76c80862be8c195e

Observation 1a55b5cc-3704-4caf-94bd-18814d579698 · outbound

This paper cites MetaOpenFOAM: an LLM-based multi-agent framework for CFD.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards MetaOpenFOAM: an LLM-based multi-agent framework for CFD

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:9cb7762ddd6b43ae6a47db14aa76a4f60b10da62dedda07979d35acbfa994e2d

Observation 2d340ce7-621f-481a-968c-f15309f5d76b · outbound

This paper cites Mooseagent: A llm based multi-agent framework for automating moose simulation, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Mooseagent: A llm based multi-agent framework for automating moose simulation, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:01c6cac65ffb0416c8ae63affe0a8cba52a9328666ae665787ebd77e44cd1372

Observation e38baeff-07f0-4af1-800e-ec6d11f182e2 · outbound

This paper cites Chronollm: customizing language models for physics-based simulation code generation.Multibody System Dynamics, Feb 2026.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Chronollm: customizing language models for physics-based simulation code generation.Multibody System Dynamics, Feb 2026

Reference 31

Resolution
verified exact
doi, observed 2026-07-14T11:30:25.748518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:33c2bfd88692ab47dc0c89733b305fc45f9539849e7adc91e39653d84413c532

Observation 913d1ec4-cdfa-4feb-a229-e36b065e6a8d · outbound

This paper cites Lang- pinn: From language to physics-informed neural networks via a multi-agent framework.arXiv preprint arXiv:2510.05158, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Lang- pinn: From language to physics-informed neural networks via a multi-agent framework.arXiv preprint arXiv:2510.05158, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:b0630f8723fee0336a5fffb37c2a66614bf8d42087563ec337ce05b459cf61a5

Observation fbfc218b-0463-4530-b5e1-d51e4e59d3fd · outbound

This paper cites FEABench: Evaluating Language Models on Multiphysics Reasoning Ability.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ca3f0805c04f979e41d8ab77c953722087c215c2c98f599ae8e5cfb937e8e9ac

Observation 72c4e00f-6067-4c45-a5a7-e4c55ce47623 · outbound

This paper cites Solving Physics Olympiad via Reinforcement Learning on Physics Simulators.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:870d7fc1a3c443250a70cb31f0e9cfce20449012741e5449a682af644d2b1cca

Observation 1359ee21-87e0-4e85-985e-3ec70721d6e9 · outbound

This paper cites PDE-Controller: LLMs for Autoformalization and Reasoning of PDEs.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards PDE-Controller: LLMs for Autoformalization and Reasoning of PDEs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ad9b1ce51153470bec5bd79cc4ec02a194ea5c4731efc66c5a11a21598c1df01

Observation 3da208fc-ebdd-4c99-9584-79ae27d54505 · outbound

This paper cites Agentic scientific simulation: Execution-grounded model construction and reconstruction.arXiv preprint arXiv:2603.00214, 2026.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Agentic scientific simulation: Execution-grounded model construction and reconstruction.arXiv preprint arXiv:2603.00214, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:18061f6e863cf4ed4c4fffd4e082b737061ca04c64c7f9f14ae36d689c0dc6d3

Observation 84a015bf-0af4-44d6-80c1-f7c88b8bed10 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards DAPO: An open-source LLM reinforcement learning system at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:c67d2c08e37113564ea04bff929f8658fbbfc26ca5ed56c2f39640489b6856cd

Observation 52864811-cc87-4474-adc2-93ba935dfba9 · outbound

This paper cites Group Sequence Policy Optimization.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f34d04500a962e8a61bf3984bcd93632fa2727b7ce1d6760ed536db3846677ee

Observation 41ec3765-e6e3-48e6-b353-8aff44293cc7 · outbound

This paper cites Stepcoder: improving code generation with reinforcement learning from compiler feedback.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Stepcoder: improving code generation with reinforcement learning from compiler feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:c454768d9cffd49ff49b964ce0d7e69b92cc5e76d01b32edecf4da7cccfac60c

Observation 082a7f23-a38a-4140-b7dd-43e98d831d0b · outbound

This paper cites Reinforcement Learning for Machine Learning Engineering Agents.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Reinforcement Learning for Machine Learning Engineering Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3e180d0609054d2e024f47c4aa1bd9dc8572e64743e9a2c10c2b3f650f2162d8

Observation f70029e2-a09f-404e-afdb-cc00f2c50e31 · outbound

This paper cites Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:d444ee7ad409615532ee16fe41b4c0b644daf8204617bd6c19444ea5207b615b

Observation 4952e689-6e5e-454c-88ea-06e30e533702 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:80242a8b9db02d259431b25d44df651eb4ef5746b54cfae400a88f158da3833d

Observation 1ec75db1-7a38-471b-b941-8b0b4239a2a6 · outbound

This paper cites Fourier Neural Operator for Parametric Partial Differential Equations.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Fourier Neural Operator for Parametric Partial Differential Equations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:458c3f25cb9fd25b7d37395ebbd2c822fcdcffc9f5d68eca660c0e59031b42ce

Observation 875278d0-a8fd-4ea5-ac02-187f5fd731bf · outbound

This paper cites Learning nonlinear operators via deeponet based on the universal approximation theorem of operators.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Learning nonlinear operators via deeponet based on the universal approximation theorem of operators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:63093609fbd93cf52d186968d91ea17edd82e8ed8822fae7d1e8a89e3615bdbd

Observation 3a8f95d5-ad38-4ff7-b096-65755577ef78 · outbound

This paper cites Towards long rollout of neural operators with local attention and flow matching-inspired correction: An example in frontal polymerization pdes.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Towards long rollout of neural operators with local attention and flow matching-inspired correction: An example in frontal polymerization pdes

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3e21a460ad0caff6b8e33115d66c6bd89a1e7ce63f19aba8174fad49262795dd

Observation 18612c69-901d-48da-853e-be40c0411fc6 · outbound

This paper cites Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:17eb16c2a67b18bffb464ed7783746eee99957968f6d989fea1f785f43373b7a

Observation c1e53791-f6a6-4825-9da6-01720d5d4ae0 · outbound

This paper cites Diffusionpde: Generative pde-solving under partial observation.Advances in Neural Information Processing Systems, 37: 130291–130323, 2024.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Diffusionpde: Generative pde-solving under partial observation.Advances in Neural Information Processing Systems, 37: 130291–130323, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:11ddee6963f958a26a128a79a702298f0d46925d454881d4c208b614f20eb255

Observation 546f6ddc-d397-4231-b52c-47b2f65662d5 · outbound

This paper cites Physics-Informed Diffusion Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Physics-Informed Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:494ccebb2894974eab7ee6c2380afffcc4f7992c2feded70080f494d2efa6ff2

Observation 0e93baef-bf6f-4707-b999-08b1e1e67819 · outbound

This paper cites Physics-constrained flow matching: Sampling generative models with hard constraints.arXiv preprint arXiv:2506.04171, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Physics-constrained flow matching: Sampling generative models with hard constraints.arXiv preprint arXiv:2506.04171, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ec444f509c42aaee524f17076281e5b952f872f8a6e675a763850051c6362b5b

Observation a1478eb7-a62d-4f7d-bbaf-00b1f33f3939 · outbound

This paper cites End-to-end probabilistic framework for learning with hard constraints.arXiv preprint arXiv:2506.07003, 2025.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards End-to-end probabilistic framework for learning with hard constraints.arXiv preprint arXiv:2506.07003, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:747553bdfb076ac43eeb1fde380aa54d0323514b57a5c6a74355988884a5516a

Observation c106f845-3b5f-4fe6-b0d8-551656f32fbe · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Training language models to follow instructions with human feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:af608dc3fd48c7852928204e3e083eb21db133d5557078976b750ecf4874123c

Observation 62d554c6-900f-468b-9109-ef3307a28e3b · outbound

This paper cites Qwen2.5-Coder Technical Report.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Qwen2.5-Coder Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:b1fb641e1459a417c709dee58edefe81e604e5cfe260554cf1f57532c0e2985e

Observation 25401330-e3f5-4936-804f-83b9a6ad3b99 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards LoRA: Low-Rank Adaptation of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:9faaa4a54e42d3e0ac368cc400da67f2e0805f43c83848d75a2003b14620a3c4

Observation 18f47ed7-4f88-4f12-8d82-8395aff5f1f8 · outbound

This paper cites A model for fast computer simulation of waves in excitable media.Physica D: Nonlinear Phenomena, 49(1-2):61–70, 1991.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards A model for fast computer simulation of waves in excitable media.Physica D: Nonlinear Phenomena, 49(1-2):61–70, 1991

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:2bd649394c7c82b8a5aefdec40d4a64fa4b01081f180ee86d00f346ff0c13011

Observation 29a80986-4a9a-47be-8f06-bdae568ae946 · outbound

This paper cites Finite difference method for numerical computation of discontinuous solutions of the equations of fluid dynamics.Matematiˇ ceskij sbornik, 47(3): 271–306, 1959.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Finite difference method for numerical computation of discontinuous solutions of the equations of fluid dynamics.Matematiˇ ceskij sbornik, 47(3): 271–306, 1959

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3df7819dcc8ae64060343390406483d1dcb0d2434f96e7e2a580decdc80f7e2b

Observation 262ef592-3736-4460-b7eb-b996676d64bf · outbound

This paper cites Towards the ultimate conservative difference scheme.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Towards the ultimate conservative difference scheme

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:3febef61f0c471c9353747116cc6ae71e166e9d3d51ccca8134053d805cc7e32

Observation b70b2f42-e862-4aeb-bc88-0fb87d62d5af · outbound

This paper cites New high-resolution central schemes for nonlinear conservation laws and convection–diffusion equations.Journal of computational physics, 160 (1):241–282, 2000.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards New high-resolution central schemes for nonlinear conservation laws and convection–diffusion equations.Journal of computational physics, 160 (1):241–282, 2000

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f2da83b9567810c624b03d27eb7562a77034182a11de5cd56d23665871b5f33d

Observation edc2e5e0-c850-4b61-8191-2e4d0b06feda · outbound

This paper cites Toro.Riemann Solvers and Numerical Methods for Fluid Dynamics.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Toro.Riemann Solvers and Numerical Methods for Fluid Dynamics

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:133a08854e1ae618df4503d22907e2d38d0b62e92d7bb01b5eb6111a8704af02

Observation 47019da4-0458-4c15-9f9a-eb857fe302e9 · outbound

This paper cites Strong stability-preserving high-order time discretization methods.SIAM review, 43(1):89–112, 2001.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Strong stability-preserving high-order time discretization methods.SIAM review, 43(1):89–112, 2001

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:5fa6878bfd87905500bc5101221540dbff3cdfb855106dc77b60cad106c03d6b

Observation 8a4015b7-9d4c-4cd6-b243-f6e25b2ba310 · outbound

This paper cites On the construction and comparison of difference schemes.SIAM journal on numerical analysis, 5(3):506–517, 1968.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards On the construction and comparison of difference schemes.SIAM journal on numerical analysis, 5(3):506–517, 1968

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:96df786251ff4eda4162a968935d4480ff06c666fee699103e83d2c8798eb0c6

Observation a52ddfc9-a4ff-4487-b268-7fc131ab12d8 · outbound

This paper cites Implicit-explicit runge-kutta methods for time-dependent partial differential equations.Applied Numerical Mathematics, 25(2-3): 151–167, 1997.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Implicit-explicit runge-kutta methods for time-dependent partial differential equations.Applied Numerical Mathematics, 25(2-3): 151–167, 1997

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:349c132aa224e4fddd48dd34d0626166d9ba10f1f261ce6e902e893f6c43d4c2

Observation 468f879b-6e28-49e1-811b-476cd6f91fd3 · outbound

This paper cites Fourth-order time-stepping for stiff pdes.SIAM Journal on Scientific Computing, 26(4):1214–1233, 2005.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Fourth-order time-stepping for stiff pdes.SIAM Journal on Scientific Computing, 26(4):1214–1233, 2005

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f2b70a43c7b3452d635ca7930ae69c91f071e1a70c61627a897226980ad7e7f7

Observation a0eeb148-5fcc-417f-8072-fb4fafdb1ba3 · outbound

This paper cites The numerical solution of the Navier–Stokes equations for an incom- pressible fluid.Bulletin of the American Mathematical Society, 73(6):928–931, 1967.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards The numerical solution of the Navier–Stokes equations for an incom- pressible fluid.Bulletin of the American Mathematical Society, 73(6):928–931, 1967

Reference 63

Resolution
verified exact
doi, observed 2026-07-14T11:30:25.737831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:48e6fdeebb25d8bb85318f0be26e78d34b0b6f011849e74e6cbf440b55a75a95

Observation c68f4517-646a-46cd-b18f-0f38ebd31974 · outbound

This paper cites Wellesley-Cambridge Press, 1986.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Wellesley-Cambridge Press, 1986

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:d026403410a4e6e35da12f500b5096e1695c47798b1f61da04b0d092dac23620

Observation 4c58c301-050a-439f-a395-ad8ceb47c94f · outbound

This paper cites Scipy 1.0: fundamental algorithms for scientific computing in python.Nature methods, 17(3):261–272, 2020.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Scipy 1.0: fundamental algorithms for scientific computing in python.Nature methods, 17(3):261–272, 2020

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:2788cb39c32a5a0ab8890982781469cf61702ae7b223012efbfa0e45444a164a

Observation ffb451eb-eecf-46eb-b60e-701e506d4255 · outbound

This paper cites On the elimination of aliasing in finite-difference schemes by filtering high-wavenumber components.Journal of Atmospheric Sciences, 28(6):1074–1074, 1971.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards On the elimination of aliasing in finite-difference schemes by filtering high-wavenumber components.Journal of Atmospheric Sciences, 28(6):1074–1074, 1971

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f31fb49d836ce9c95c59abc4f5d6eadd71116463ae17653ea2941b1f8b71929d

Observation 68bbd1b8-127d-4493-8bbb-b0d3b0c5a434 · outbound

This paper cites Semi-lagrangian integration schemes for atmospheric models—a review.Monthly weather review, 119(9):2206–2223, 1991.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Semi-lagrangian integration schemes for atmospheric models—a review.Monthly weather review, 119(9):2206–2223, 1991

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:98f9871b58bec049a91fc2828b99344f8d025f6a9121986728af440f69f1aca7

Observation 4458a449-2ca6-4396-9a45-40c8850cd024 · outbound

This paper cites SIAM, 2007.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards SIAM, 2007

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:7a879c1f6558225ec98f612bf27ec99ca86fa3dc9b66ebfb9eb8df19af4f66cc

Observation cccf25a3-36ab-4c17-9633-25e48b97dfd3 · outbound

This paper cites The calculation of the interaction of non-stationary shock waves and obstacles.USSR Computational Mathematics and Mathematical Physics, 1(2):304–320, 1962.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards The calculation of the interaction of non-stationary shock waves and obstacles.USSR Computational Mathematics and Mathematical Physics, 1(2):304–320, 1962

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:2e0bb3522807db5e0166625bc5cd536f5f9fcd03c1ec32919cce952915e03b45

Observation 2826c273-9ea2-42ed-b809-f5e36decb625 · outbound

This paper cites Systems of conservation laws.Communications on Pure and Applied Mathematics, 13:217–237, 1960.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Systems of conservation laws.Communications on Pure and Applied Mathematics, 13:217–237, 1960

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:9313a94bd5266a0e3344b19b70a79864074f91035c12ee1cebc5e73ae0e72a9c

Observation f3bceee2-ffed-4c61-98dc-29d9215b16fb · outbound

This paper cites The effect of viscosity in hypervelocity impact cratering.Journal of spacecraft and rockets, 40(5):757–763, 2003.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards The effect of viscosity in hypervelocity impact cratering.Journal of spacecraft and rockets, 40(5):757–763, 2003

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:28335f79d7f1ea495806d5bdccf957665b3e50e98ee8ef84414ff6520212291c

Observation df339ee1-c52c-457b-bd41-89605c43c313 · outbound

This paper cites Springer, 2003.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Springer, 2003

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:d3451416f1801340617d23829d1ad67b29603e7f9aa4c4db375b3e4aaa44557d

Observation d93d8694-048c-474b-b93d-9c9e2cdedd53 · outbound

This paper cites Finite volume methods.Handbook of numerical analysis, 7:713–1018, 2000.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Finite volume methods.Handbook of numerical analysis, 7:713–1018, 2000

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:b647eb52cae679ba3992dc13d48b8172e895ad93c90d15390e9e855a9183f66d

Observation cdda3a8f-19cc-4c8d-b1da-891fd4a6ac2e · outbound

This paper cites Methods of conjugate gradients for solving linear systems.Journal of research of the National Bureau of Standards, 49(6):409–436, 1952.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Methods of conjugate gradients for solving linear systems.Journal of research of the National Bureau of Standards, 49(6):409–436, 1952

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:fa67c938050728316ccfbe89db2bc992e696abb012c5e8b3df3d6caf2fd7db0f

Observation ef4648b6-41a3-46da-8d2e-b9292b759f08 · outbound

This paper cites SIAM, 2003.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards SIAM, 2003

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:9023728e829090700b584e4c4f45f5406aa3751eb676db4f8cb32b00db41930a

Observation f246ce2d-75ae-44a4-b0d5-9da9de767a50 · outbound

This paper cites Springer, 4 edition, 2020.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Springer, 4 edition, 2020

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:e24b08870fe1b2db3263d0135f3a9e5f1e145d50bccf7bda7c196572a8da15b4

Observation 3ae2cb35-29a7-43c5-9805-e18c08913990 · outbound

This paper cites Approximate riemann solvers, parameter vectors, and difference schemes.Journal of computational physics, 43(2):357–372, 1981.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Approximate riemann solvers, parameter vectors, and difference schemes.Journal of computational physics, 43(2):357–372, 1981

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:a50ccc64514364985e2ee27b965cc4511d81f773d0a967583caff7d974404c19

Observation eb47f0e1-c1c6-4c64-b552-c5e5c29897be · outbound

This paper cites Information theory and statistical mechanics.Physical review, 106(4):620, 1957.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Information theory and statistical mechanics.Physical review, 106(4):620, 1957

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:f7bd8c5c55c3bcaa510b6a0d1e9a67067c205cfa70c7fdfd0306726c6659d806

Observation 48dbb034-9c96-4bb9-b7f8-c22117313ee6 · outbound

This paper cites Academic Press, 11 edition, 2014.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Academic Press, 11 edition, 2014

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:590e85b6bb4edaa48196f540181f023314c93d45f3aa1e57e47638bdc30c41a8

Observation 55b2cfab-85c7-4f2e-ac09-efe89111e1e0 · outbound

This paper cites Equation of state calculations by fast computing machines.The journal of chemical physics, 21(6):1087–1092, 1953.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Equation of state calculations by fast computing machines.The journal of chemical physics, 21(6):1087–1092, 1953

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:85e16cefee76c9e30e15115f14281f5afec97f4db9789b6f01776e0675382e1a

Observation 8877ab60-e240-4a88-82e8-1fbcb8052472 · outbound

This paper cites Ziebart, Andrew Maas, J.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Ziebart, Andrew Maas, J

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:43fe79fbaf9cee25d284950795645da2b6bc9ce2cfff3130efa4ab14c44964fe

Observation 8cf82fa4-923f-4b76-b026-9384fe8b1903 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:2ac78005ed05a5b2239d8cfa2c376aed83b5e4e7ea0ecede735be99ea2444964

Observation 6d26fc3b-1ddb-4815-8dc5-f8858144caac · outbound

This paper cites Where applicable, we additionally require that solvers within a scheme family agree on smooth initial data to within their formal order at fixed resolution.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Where applicable, we additionally require that solvers within a scheme family agree on smooth initial data to within their formal order at fixed resolution

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:c8e57791a61878907d82da32b1e8440db264c0fcf6827de0bc8ced3ac7614124

Observation 499fcfd1-cbfd-4bbf-99c1-f1e06c0457ae · outbound

This paper cites This producesabsoluteerror against the exact ue at each grid level rather than a self-consistency ratio, and detects sign errors and stencil bugs that self- convergence cannot.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards This producesabsoluteerror against the exact ue at each grid level rather than a self-consistency ratio, and detects sign errors and stencil bugs that self- convergence cannot

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:95e22688fff303125510b832c7df159972061e9cd066dd1d4b4981b905e205a4

Observation c849d7e5-efa8-43c8-9412-8a447c9a0809 · outbound

This paper cites This provides a check against an independent benchmark, complementing self-convergence and MMS.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards This provides a check against an independent benchmark, complementing self-convergence and MMS

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:10a6f3ec14952767b04b2c6b5ddce6d003b57875413fed253ff672ab592d08a8

Pith citing papers

No inbound Pith citation observations are available.