Pith. sign in

Paper Citation Record · LEDGER

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.29185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29185 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:09:09.970266Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dc2264c-ba80-4fbc-a6a6-4e29844905e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.872221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.872221Z digest=sha256:a3351ade33b290d8a676a6b43b67cb9ba5a9f33ea822420396846e11dfc1a0ae

Observation 41d2d2aa-425d-4996-86ab-e048a4a8edbe · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.997789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.997789Z digest=sha256:1d771a12044b3f81e3aa79a35cb6d3c3e96d2d25ff043163a7eb3e6d5b646661

Observation 431449d6-8799-4519-83f6-9f9298c0e0a0 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.149160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.149160Z digest=sha256:ac2824aa2355f3493766280925b24c80e579162b6c1f392203cb48f5e67922e0

Observation 98425e2c-411b-4603-ad37-b60a5b6deac3 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.218493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.218493Z digest=sha256:ec4d9805f3c8d27e4b89f42d5e4e96b7426570bd6baacbe75a801f65e426f200

Observation 253f7795-9355-42b0-a36b-e9ecf7a3f2e7 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.366566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.366566Z digest=sha256:ef8f7e33308e5c6cfd40719aa5ca64506d3b108e2808fe58793d9a68906ae3af

Observation 999280f8-9d1a-41e4-b281-f26268dbd9c0 · outbound

This paper cites 2026 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , url=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.482182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.482182Z digest=sha256:ed3136cb954885243495235df8062baa1b3771a241f4783e8a63bac2158f2349

Observation 5c0d7c3c-cbfa-4876-8a0a-40a36dd1f089 · outbound

This paper cites J1: Incentivizing Thinking in.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End J1: Incentivizing Thinking in

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.597246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.597246Z digest=sha256:d6b8a498dd9e68435b9cb92f5f34abfbe7df366efef0dca51dad92a47835f7e6

Observation fc97e75a-547a-4643-a6e4-d3e1d554138b · outbound

This paper cites Learning Structured Output Representation using Deep Conditional Generative Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Learning Structured Output Representation using Deep Conditional Generative Models , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.676252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.676252Z digest=sha256:ff0f887126476295a437cb8626a80e310002beaf4a1c8c2d05caeda1fe01929e

Observation bbfaef32-3e3b-49d4-95dc-0a9b9cb743b8 · outbound

This paper cites Training language models to follow instructions with human feedback , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training language models to follow instructions with human feedback , url =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.762707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.762707Z digest=sha256:033f60e4228a2dd71514a383debb6954f15b09ae386f936348de2476329df155

Observation bb2aeee2-e7ee-4f98-b88c-ac1539f2f3e4 · outbound

This paper cites InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.944216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.944216Z digest=sha256:d3517cbc13153bbfd5aa9dcbfac5e0e9d038a3bd7d736b54643227fb191326c8

Observation 2c49e1ae-d586-40e7-958e-1338a92d86f5 · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Improving Reward Models with Synthetic Critiques

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.175902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.175902Z digest=sha256:a196e15b81767cf3a76816ab3238960a1e3962b4759eea169fc51edcccf10cb1

Observation 80d06d8e-4780-42aa-895d-3a4d513f0c70 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:09:05.271074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.271074Z digest=sha256:42a1eea9bb1becda7e2e2d3e634947077f4de761e2d55ea7149a2a35ececf0fe

Observation 485eac96-9550-4488-868a-acade5e44020 · outbound

This paper cites Critique-out-Loud Reward Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Critique-out-Loud Reward Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.397076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.397076Z digest=sha256:82506ece127e51421fd7a7e26998b0d16e2f30d45634a84540da3f625f0db2b8

Observation 98922cab-3159-4e19-8590-cf66749076a7 · outbound

This paper cites Generative Judge for Evaluating Alignment , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Generative Judge for Evaluating Alignment , url =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.493941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.493941Z digest=sha256:f12e6c7a152ede591a8dea9b066e14ae451fabdf63ddfb4ac6a6806c05d5e8b9

Observation 4f9561c5-399d-492f-be3c-31982f207d63 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Deep Reinforcement Learning from Human Preferences , url =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.653095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.653095Z digest=sha256:13ffd461c541844ac6f7da662513b0efdfefcd78d9c384f0981b41ebd61280d1

Observation c9c45d56-12f8-4fa6-9f17-2e48de8854ed · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Twelfth International Conference on Learning Representations , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.784919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.784919Z digest=sha256:ed66849760f8b69b3b7105cd9d3920158c430df61fe25c0eb2a45ae9fc73850a

Observation c87ef72f-a19a-4448-b63a-45012dbb70cc · outbound

This paper cites Terry , journal =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Terry , journal =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.923413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.923413Z digest=sha256:bb35ade7fbc6213f84ab83e1455d9c908cf6fbb3b010e17f3ef124ccae3ac4df

Observation 99d8a67f-44b0-43b2-941a-ae5ef0deb5b8 · outbound

This paper cites 1959 , publisher=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 1959 , publisher=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.009105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.009105Z digest=sha256:165565e7bbb0d8711617646a13570559975699e10abacf68c13b0cbe433fa187

Observation b7c64419-c3b0-4a5b-a4b3-48c316a6f2af · outbound

This paper cites Rationalizing Neural Predictions.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Rationalizing Neural Predictions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.107054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.107054Z digest=sha256:2a1c85efab03b4baec78f20c4359314e857910302592f3bb813e75bc6e07f052

Observation 73f6cb0e-d506-4901-a314-8e71d4b2359c · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End STaR: Bootstrapping Reasoning With Reasoning , url =

Reference 20

Resolution
verified exact
doi, observed 2026-08-03T12:13:46.150492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.296072Z digest=sha256:4d207925e862e9c6177275bf3a8cc6d5f3d467358c3c9ef1f3411f42563a6898

Observation be55845e-65b9-49d6-be88-4b0e8f15c7b4 · outbound

This paper cites Saurous, Rif , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Saurous, Rif , booktitle =

Reference 21

Resolution
verified exact
doi, observed 2026-08-03T12:13:45.934005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.404409Z digest=sha256:cdee4630f63d3724d3861dfa61c5d3acfc1bb7fb16f03d4ad16d76ffdb99b56a

Observation 58a2f00e-2cec-4848-a7b3-052e9fb2700e · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.520255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.520255Z digest=sha256:6846ad82a57adc6c88094e4dd561421e9ea006dbd255c60830d4ab9df42762db

Observation 8c4cbae2-b25b-4815-ac02-52a2e1872a01 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.747386Z digest=sha256:dbf37279b3a73fb315a46460c83599d421ae17ab33c6988c34dbb475e43f2ef0

Observation 886a19f1-4caa-44b8-b525-559dbd02ae8e · outbound

This paper cites Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.931488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.931488Z digest=sha256:3cbb2d9b6d0c11e7b555d75976a2e6b0e38a6559146b3affd95dd14366238397

Observation a255d2be-95fb-4824-8bac-d838c86dd918 · outbound

This paper cites 2024 , editor =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , editor =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.125029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.125029Z digest=sha256:38338dccba8700d5d6a773b4eb0cd87cefe54687cd7e68f4be681fad1d8b2d35

Observation 7d56f8fe-1f77-469b-b8b3-0e315a405ce6 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.287104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.287104Z digest=sha256:f698c95ddf23151d4645ebcaf7289196d607683ae45ba4cc51ad96153a8f0b65

Observation 7c3a15a3-b696-4f7f-bf85-4372992b453a · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.449364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.449364Z digest=sha256:a7cdc311e6ad11083e16ee93f4d5c58eea02ef9a62aab7daeeacd054515e4040

Observation 174e1aaa-f415-415a-bc21-aaa61d76f3f3 · outbound

This paper cites WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.583260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.583260Z digest=sha256:d02a617e8acf22cb969cc3dd1af4f1afe16c30b83578094239d8fe647da14e8a

Observation b64833c2-6236-48f6-aca4-01d4997b3798 · outbound

This paper cites O ffset B ias: Leveraging Debiased Data for Tuning Evaluators.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End O ffset B ias: Leveraging Debiased Data for Tuning Evaluators

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.793737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.793737Z digest=sha256:9b5e8efb9654a44d179771beddd0e48b37867fba0fe5aaf5e9dc2bd0238fa0ae

Observation 293f4311-98c8-4572-b693-47dc9ffd176a · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Forty-second International Conference on Machine Learning , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.896547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.896547Z digest=sha256:f26ecc0609ff5dfeb655943e83dd847840f29cf55afb8cf9c32f882455488723

Observation 90969518-4203-4f4d-984b-a76806aaa611 · outbound

This paper cites and Lee, Jason.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Lee, Jason

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.999068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.999068Z digest=sha256:84cdd9ee5f4cf606045c89fc17b36e62ca6369b7dd36f341be7d1172a6e0430d

Observation 4302bb35-3fe1-4775-a68d-d3db24b1de73 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.110251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.110251Z digest=sha256:b5c5565b2f675cbe04e6593c0086f909afd9b387f64db849b8e40c026685947e

Observation c18b9ce1-29b8-457f-af93-eddc64eb3203 · outbound

This paper cites and Hajishirzi, Hannaneh.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Hajishirzi, Hannaneh

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.182976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.182976Z digest=sha256:3a7250cdea2925638528e2e310e37066ed85e61f4a8c0c86ed6aeaa3769abb67

Observation b667365d-1ffb-4911-b7b3-517dfc6962aa · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Fourteenth International Conference on Learning Representations , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.277734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.277734Z digest=sha256:e69041732458b996c81e9c500af4626e992c0350eba1fd9160d5bd63494c9c81

Observation fed329e1-8fd5-4f2f-ac92-9af408c9a6ff · outbound

This paper cites 2025 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , url=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.375921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.375921Z digest=sha256:4796a7937d5826536a2360c1bf090ccbf2acdb9079ce19655997552741d7483e

Observation 79867b3e-eb47-433e-b966-e7bbc593a997 · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Gonzalez and Ion Stoica , booktitle=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.484852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.484852Z digest=sha256:26bf73c7271c556d8ec0e7a097aa4bb3f6a2d1881e54a21663d145aeccbfb5d7

Observation 8b8fa87e-aa81-4eac-a9fe-eaf6ac4e9f53 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End JudgeBench: A Benchmark for Evaluating

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.643164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.643164Z digest=sha256:c6ab2fdacc45eae35572e61c44c2e27d35bbb93b98e38905420bfdcf974176e6

Observation 218e1456-c11e-4cdd-8322-a84a0339dcef · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.759591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.759591Z digest=sha256:0c3660173bdd5713be34d07341649ad3fedc69718532c14cb739c8c559f6b5de

Observation fc2aeba1-47cd-4469-acff-8e1c2b21b2e4 · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.853125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.853125Z digest=sha256:0f6c5abc98d5406bfea836d0e0bec8a0098360d4ac3562304fc4e017412a4f61

Observation 0076fc1e-1ad7-4427-aa0b-1a04f24cd55c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.946082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.946082Z digest=sha256:80361d7eb239fd82b5f351e97c213f627ededda0d10ecb45cb29225aaf0b876c

Observation 87186d91-407e-45f9-9956-6060da303d7c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.065374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.065374Z digest=sha256:60fc62727e5230ddb177323ec7c667272060c38812139aa92ffa3fc4829b5cb1

Observation fa4bf8ca-2535-4ee2-8908-0f9ab5e3a769 · outbound

This paper cites Second Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Second Conference on Language Modeling , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.185899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.185899Z digest=sha256:d3d93a10661ceca80d7837e81f7974ec2444c81fa8d25e648767dc44ae31fbce

Observation 1c6d41ef-8c97-46da-925a-a457b1746d47 · outbound

This paper cites 2024 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.348666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.348666Z digest=sha256:7cb164bac8c30495a70400e9654cd19c019abfd3ccec798db9e6116d723ecdf8

Observation 78c659ae-7380-4c7b-bb48-9f6a876a6875 · outbound

This paper cites First Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End First Conference on Language Modeling , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.492302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.492302Z digest=sha256:20747740ce41795ee15be8d3e6a76965af3aefcf97d3a9ffb8c330d68cdc0b04

Observation 8ccac31c-3c99-483c-a628-34cdebbdedb8 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.677129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.677129Z digest=sha256:3448eb9e4ab1030400026fbd6f9d76e89f8ba4614deff06b66c7d96811a696d5

Observation cef30655-655a-4144-9f6b-2b1535ee6cc5 · outbound

This paper cites and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.790291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.790291Z digest=sha256:1edbe2e94d732f5e40ca71800226646f964086f48746482797eb0582a672629a

Observation f2de6b5f-bd72-4163-810d-d55b1e16c122 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.970266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.970266Z digest=sha256:99af5bff342bc77a620163d1d3d186c5f41a14bfc474b701bb811e0df8f93396

Pith citing papers

No inbound Pith citation observations are available.