Pith. sign in

Paper Citation Record · LEDGER

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.29185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29185 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:09:09.970266Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dc2264c-ba80-4fbc-a6a6-4e29844905e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.872221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.872221Z digest=sha256:01f4ede018a3267913f9a8a64e3fcbf5c5f31691dbe655e1c3c86fbc12ca52f1

Observation 41d2d2aa-425d-4996-86ab-e048a4a8edbe · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.997789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.997789Z digest=sha256:64b07aaa69c695459a91c742809c6bc12d01f7a35856fa23edbebe4a038c4141

Observation 431449d6-8799-4519-83f6-9f9298c0e0a0 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.149160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.149160Z digest=sha256:da3c9b5871211df12f6cd0701ff6f36ec852db078f684869a0c94649984b205e

Observation 98425e2c-411b-4603-ad37-b60a5b6deac3 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.218493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.218493Z digest=sha256:e74d012357b7d41198fd2676a307d7992de39e1138cfd6889f9369ff0dda6579

Observation 253f7795-9355-42b0-a36b-e9ecf7a3f2e7 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.366566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.366566Z digest=sha256:3b749e9133360b9629841195d7475dd8a5f6b316f75eaa964a8a6aa82ec6d47f

Observation 999280f8-9d1a-41e4-b281-f26268dbd9c0 · outbound

This paper cites 2026 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , url=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.482182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.482182Z digest=sha256:0e407fefa055e454af0b00f835f6b57fe112c93319f7617e280ab2b05c1816dc

Observation 5c0d7c3c-cbfa-4876-8a0a-40a36dd1f089 · outbound

This paper cites J1: Incentivizing Thinking in.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End J1: Incentivizing Thinking in

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.597246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.597246Z digest=sha256:06007c3dbe7e7d841e9da9b103dba310fd10f2514199d7385595e8d1c80808bd

Observation fc97e75a-547a-4643-a6e4-d3e1d554138b · outbound

This paper cites Learning Structured Output Representation using Deep Conditional Generative Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Learning Structured Output Representation using Deep Conditional Generative Models , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.676252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.676252Z digest=sha256:d6de0bff619e57dca68e25c476627fa1e8fe039871b0bae21973078fef95d0d8

Observation bbfaef32-3e3b-49d4-95dc-0a9b9cb743b8 · outbound

This paper cites Training language models to follow instructions with human feedback , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training language models to follow instructions with human feedback , url =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.762707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.762707Z digest=sha256:4b6b9437563fc888db8b3f349ee50856f1d7e0a2e0433ce615cd1dc299977329

Observation bb2aeee2-e7ee-4f98-b88c-ac1539f2f3e4 · outbound

This paper cites InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.944216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.944216Z digest=sha256:ebd1dc07a9663cd96a0561884976f963bac8f1a6d167e19818e185dbf1a43121

Observation 2c49e1ae-d586-40e7-958e-1338a92d86f5 · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Improving Reward Models with Synthetic Critiques

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.175902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.175902Z digest=sha256:3ad86d5ec98e4279081cac5b317224a73726b16109b108d280796ee397fcaaf0

Observation 80d06d8e-4780-42aa-895d-3a4d513f0c70 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:09:05.271074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.271074Z digest=sha256:b02f8c1a1b283d85ba05ce6208f4a8766e52325ba42b575bccb9268d314c8ee2

Observation 485eac96-9550-4488-868a-acade5e44020 · outbound

This paper cites Critique-out-Loud Reward Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Critique-out-Loud Reward Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.397076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.397076Z digest=sha256:c3c91e677dedfe087a34821c1c568b927ed77a2992f05c3931d0eda73f4c06e1

Observation 98922cab-3159-4e19-8590-cf66749076a7 · outbound

This paper cites Generative Judge for Evaluating Alignment , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Generative Judge for Evaluating Alignment , url =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.493941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.493941Z digest=sha256:f5be51cc3492e8b875aa0a5a2d4cac03b8e7cac270a2e4b0cba30e70728317f7

Observation 4f9561c5-399d-492f-be3c-31982f207d63 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Deep Reinforcement Learning from Human Preferences , url =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.653095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.653095Z digest=sha256:49d351a41ee1cf44c6899abc94c0ce91d2a7f85ca865b76e66fb625c9bed15a1

Observation c9c45d56-12f8-4fa6-9f17-2e48de8854ed · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Twelfth International Conference on Learning Representations , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.784919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.784919Z digest=sha256:52a26c8d71e9205895d38d9caeba837822143e816581defd705b59aa0f941ae8

Observation c87ef72f-a19a-4448-b63a-45012dbb70cc · outbound

This paper cites Terry , journal =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Terry , journal =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.923413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.923413Z digest=sha256:acd0885c66a137ecd3aa26e2459bb56e701c6788dc69f76050442c556aef5a15

Observation 99d8a67f-44b0-43b2-941a-ae5ef0deb5b8 · outbound

This paper cites 1959 , publisher=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 1959 , publisher=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.009105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.009105Z digest=sha256:659147815981f5f445a7f2b258c7671816a84fdff1ba97635028e7a12bf576b9

Observation b7c64419-c3b0-4a5b-a4b3-48c316a6f2af · outbound

This paper cites Rationalizing Neural Predictions.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Rationalizing Neural Predictions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.107054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.107054Z digest=sha256:d5bf25da78a7ee9f0fba721c41157ed77511163d8e3eb1c8abd3f9371d1ee332

Observation 73f6cb0e-d506-4901-a314-8e71d4b2359c · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End STaR: Bootstrapping Reasoning With Reasoning , url =

Reference 20

Resolution
verified exact
doi, observed 2026-08-03T12:13:46.150492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.296072Z digest=sha256:e668ca00f55c2d52873e12ea0875db8d664379960ccfae46589e2d66ffde1c63

Observation be55845e-65b9-49d6-be88-4b0e8f15c7b4 · outbound

This paper cites Saurous, Rif , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Saurous, Rif , booktitle =

Reference 21

Resolution
verified exact
doi, observed 2026-08-03T12:13:45.934005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.404409Z digest=sha256:5a57fb4353746d31de0ae042d9d89fd427c5712ff906181edae0fefd63b4abfa

Observation 58a2f00e-2cec-4848-a7b3-052e9fb2700e · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.520255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.520255Z digest=sha256:42fe222197e280060edc5f06dac7a2eee14620587512513967f23fd817bf8825

Observation 8c4cbae2-b25b-4815-ac02-52a2e1872a01 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.747386Z digest=sha256:cb056e76def31feccbe6942dee5f6e6752438705d23e495b24dae974d5e200d2

Observation 886a19f1-4caa-44b8-b525-559dbd02ae8e · outbound

This paper cites Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.931488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.931488Z digest=sha256:69eaff1b60353a1bb537e140a68c77b1aed15074032a1d7ea80a3784d3cbc2b7

Observation a255d2be-95fb-4824-8bac-d838c86dd918 · outbound

This paper cites 2024 , editor =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , editor =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.125029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.125029Z digest=sha256:c46686752540ede5f9f3d0c4d1f4f3fdc1bfb74573eb6a16a3ae57a8e6af0036

Observation 7d56f8fe-1f77-469b-b8b3-0e315a405ce6 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.287104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.287104Z digest=sha256:2ab9e242801162d3ba6ff8f217d1dc97ac50c09900acb7a9fdf750ebcc4bd9e3

Observation 7c3a15a3-b696-4f7f-bf85-4372992b453a · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.449364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.449364Z digest=sha256:925af1597ef72b052ddec11aca41b24a85dd1419ad81d646c5dc76e67677f61a

Observation 174e1aaa-f415-415a-bc21-aaa61d76f3f3 · outbound

This paper cites WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.583260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.583260Z digest=sha256:a5c59ab39cb21b07a6e0fedde076b7505a48d17c1390d4291d22a91f014baa2b

Observation b64833c2-6236-48f6-aca4-01d4997b3798 · outbound

This paper cites O ffset B ias: Leveraging Debiased Data for Tuning Evaluators.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End O ffset B ias: Leveraging Debiased Data for Tuning Evaluators

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.793737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.793737Z digest=sha256:a4b5e911dd79d4e94bc515eb4bc369c200c8489fd4c4fc7f8728c81bf1881a51

Observation 293f4311-98c8-4572-b693-47dc9ffd176a · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Forty-second International Conference on Machine Learning , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.896547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.896547Z digest=sha256:57c7916812750a402b30772cc2ae87146d68bbab056bb2a540d1a916f4246e31

Observation 90969518-4203-4f4d-984b-a76806aaa611 · outbound

This paper cites and Lee, Jason.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Lee, Jason

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.999068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.999068Z digest=sha256:9c869a8fd24d8333d8fdd753ce08874a37f0a1c01dd29452a3a0fb290de7e45d

Observation 4302bb35-3fe1-4775-a68d-d3db24b1de73 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.110251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.110251Z digest=sha256:230e86614216e59a97690037966cf7f66f2565ae39e50fe7a0db452827cfabed

Observation c18b9ce1-29b8-457f-af93-eddc64eb3203 · outbound

This paper cites and Hajishirzi, Hannaneh.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Hajishirzi, Hannaneh

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.182976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.182976Z digest=sha256:5727ff28c295d50eaf601879caa03252026a1192a22fc0e8fa7571f76e840164

Observation b667365d-1ffb-4911-b7b3-517dfc6962aa · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Fourteenth International Conference on Learning Representations , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.277734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.277734Z digest=sha256:67c546ca3857e055de2131c9ae3c62a5e2fb261b1f5e5968ca24ec213ad45227

Observation fed329e1-8fd5-4f2f-ac92-9af408c9a6ff · outbound

This paper cites 2025 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , url=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.375921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.375921Z digest=sha256:65f5217e20423d19b7667936869f9f4c37ccca27c37dd504fc0216f35a0d8679

Observation 79867b3e-eb47-433e-b966-e7bbc593a997 · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Gonzalez and Ion Stoica , booktitle=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.484852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.484852Z digest=sha256:f9001e0c728fc667f10041beb2c24fa63df5d7e9335a51e436f976e01f4e9ac4

Observation 8b8fa87e-aa81-4eac-a9fe-eaf6ac4e9f53 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End JudgeBench: A Benchmark for Evaluating

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.643164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.643164Z digest=sha256:3f1b4565a7439a3fb8cbfc0bc08d85a50309ed22010ab4aa678d942f1152cd01

Observation 218e1456-c11e-4cdd-8322-a84a0339dcef · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.759591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.759591Z digest=sha256:cfebfd4129c0ed78ba5d78b53c30bdc8d9c3c70bf5ccc7086657b6af7c25f3c6

Observation fc2aeba1-47cd-4469-acff-8e1c2b21b2e4 · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.853125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.853125Z digest=sha256:0f4c07933b9f3d6f62404a0f796239ec59f5673d88a99fd1e0efa22cb51c1b1a

Observation 0076fc1e-1ad7-4427-aa0b-1a04f24cd55c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.946082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.946082Z digest=sha256:f2de32437a25f0b437175df5cde1cb6da2bbb78c6bd4e3dce40344b135e33824

Observation 87186d91-407e-45f9-9956-6060da303d7c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.065374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.065374Z digest=sha256:790364256e4ca6d64bcef2b426aada3187078486f2060ab3de255c73b32067ac

Observation fa4bf8ca-2535-4ee2-8908-0f9ab5e3a769 · outbound

This paper cites Second Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Second Conference on Language Modeling , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.185899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.185899Z digest=sha256:3b90e8b5e739dabb7d35c6dd004b218cb23fac857cefe8f0b69dd06be31b060b

Observation 1c6d41ef-8c97-46da-925a-a457b1746d47 · outbound

This paper cites 2024 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.348666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.348666Z digest=sha256:1c1d6f63a45cf32df93d7f4d00063bfac0a07dd2c19f9c298b1582d6858f802d

Observation 78c659ae-7380-4c7b-bb48-9f6a876a6875 · outbound

This paper cites First Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End First Conference on Language Modeling , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.492302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.492302Z digest=sha256:53e4a7aa33b4c9809a3e920fcc74f97b7c803329f3c96cfe8769a241dfb531a2

Observation 8ccac31c-3c99-483c-a628-34cdebbdedb8 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.677129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.677129Z digest=sha256:9a1e7299d318b12b4e2c858131b58354d8ab294ac42d10fcb78eaf7c09cce86d

Observation cef30655-655a-4144-9f6b-2b1535ee6cc5 · outbound

This paper cites and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.790291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.790291Z digest=sha256:4b4f290416465290543771cd57016ebc8be59f1557a421499ee9696aa1ff28e9

Observation f2de6b5f-bd72-4163-810d-d55b1e16c122 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.970266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.970266Z digest=sha256:d605151432e34ffbc3496cdf5c505f6c8702880bf986ce6ea942e99c51268106

Pith citing papers

No inbound Pith citation observations are available.