Pith. sign in

Paper Citation Record · LEDGER

Uncertainty-aware Reward Design Process

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.02256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02256 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:15.423614Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3993cf90-b9c6-4592-b88f-ee2c2f36d6c1 · outbound

This paper cites grasping.

Uncertainty-aware Reward Design Process grasping

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.020372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.423614Z digest=sha256:25f16ce252e85ecfc23c2ffee5538cfb7bdb2cbc89aec89d616a45341d21bc0c

Observation 701ca3bc-1972-4aff-9005-26a8fc6946d7 · outbound

This paper cites (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n.

Uncertainty-aware Reward Design Process (9) By performing coordinate scaling transformation on the sample points, we obtain new sample points p(i) = (˜p(i) 1 /l1,··· , ˜p(i) d /ld), i= 1, 2,··· ,n

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.200589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.295720Z digest=sha256:550f2e2335d625fd04825e02910fb0fb261b5b9e41aba6e0ec8cb26b732b9b66

Observation 8a56d682-ab05-4f3e-884a-5c185937ea75 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Uncertainty-aware Reward Design Process From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.937399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.937399Z digest=sha256:957a38c662b22686ced87db65e5700ce4bbb8d329d6696cfacbd338ac93ad6c9

Observation 74d66903-5814-4b75-bce6-d39ba09089d6 · outbound

This paper cites DeepSeek-V3 Technical Report.

Uncertainty-aware Reward Design Process DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.052660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.052660Z digest=sha256:74fb86477249b93924b0fad0438ffefd121b4516aeca8c022a0813204904186d

Observation 0365d674-560f-4a63-978f-d41d6024c0c5 · outbound

This paper cites GPT-4 Technical Report.

Uncertainty-aware Reward Design Process GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.232861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.232861Z digest=sha256:10a81478452d977b374d7ef824eb20622f4551dbe871f085e49d3b8c6602804e

Observation da250cde-64fe-400e-8ace-fa4420b0013b · outbound

This paper cites Qwen2.5 Technical Report.

Uncertainty-aware Reward Design Process Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.325053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.325053Z digest=sha256:0768994238655420d7760839746828b1a10284385df1b7c1c3f87d77c7292135

Observation 3a3d6b55-02b2-4f9b-b2ea-80a85786bd9f · outbound

This paper cites A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions.

Uncertainty-aware Reward Design Process A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.548517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.548517Z digest=sha256:f6bac08151e4b94000efbb6178b1742b85029a7799065efba20e3945a2ff03f9

Observation 20e09c50-e40b-4152-8b15-49e021975eb7 · outbound

This paper cites Api is enough: Conformal prediction for large language models without logit-access.

Uncertainty-aware Reward Design Process Api is enough: Conformal prediction for large language models without logit-access

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.664937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:14.711229Z digest=sha256:a56a6e628d6b5effe20de7432f993963799ff231f2e63005df8d169228b930ab

Observation 9af2285d-9255-47db-b961-bd4976b7c8c9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Uncertainty-aware Reward Design Process Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.805208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.805208Z digest=sha256:a15102e0653d2fe1d075888ada6cf554699e4bbc9728c24da30ceefca455654b

Observation bb6c000c-b41b-4496-bf78-f820ad8f56e3 · outbound

This paper cites Do phd-level llms truly grasp elementary addition? probing rule learning vs.

Uncertainty-aware Reward Design Process Do phd-level llms truly grasp elementary addition? probing rule learning vs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.898533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.898533Z digest=sha256:2aa4ccad2d13dc7464a912d910827a2bf5eb7d5b3341974263b865f6a34af909

Observation 52545019-4836-4976-b621-651dc3176642 · outbound

This paper cites Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation.

Uncertainty-aware Reward Design Process Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.005796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.005796Z digest=sha256:7d54d354ee92c8af2c055bcc0eaf1a352bcd5c0440d5e15124c45d1a083d9f32

Observation 15327041-fc96-46a8-b3bf-afecdf26e69a · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Uncertainty-aware Reward Design Process A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.150337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.150337Z digest=sha256:0d074b91365876df2dd7bfaab2625cfb10b5d1b8b9163690820cef1aec248df6

Observation 3da67d1c-7edf-438c-b3e6-07fe03250df9 · outbound

This paper cites The robot must push a movable chair from its initial location to a designated target region.

Uncertainty-aware Reward Design Process The robot must push a movable chair from its initial location to a designated target region

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.458326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:15.250787Z digest=sha256:998e4b2855979228cc567fba12dc7e628a8df0893f7d794d970bd481a1c4daaf

Observation 85bba71d-bfa2-4543-b947-d32463ec5e82 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Uncertainty-aware Reward Design Process Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.628984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.628984Z digest=sha256:0968b6a3be609c4571a4952bde1878b426b078ff80314a2445ca99f8b0dbdac5

Observation f8f18e13-f230-4c90-a7b1-85703a47297b · outbound

This paper cites A survey of confidence estimation and calibration in large language models.

Uncertainty-aware Reward Design Process A survey of confidence estimation and calibration in large language models

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:17.193725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.659017Z digest=sha256:3e55caf095d19b27e8876dd9c03a26fc27a7ef0d52e6349446c17b7fc83903a7

Observation 353afe68-687c-4b83-b7af-acde704e0299 · outbound

This paper cites Reasoning with language model is planning with world model.

Uncertainty-aware Reward Design Process Reasoning with language model is planning with world model

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:39:16.938341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.726804Z digest=sha256:766d5e7f7f0780879a95af6d430aebf81e57f0dc9eaf9fa4c38d554604e0d02b

Observation 38ee5e0b-970b-44c7-9c04-a379d987b780 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Uncertainty-aware Reward Design Process Proximal Policy Optimization Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.376825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.376825Z digest=sha256:2e869dd6c36354f5b85c181c7ef1be4222c340626cd82d5b7a9aa8a3479b64f8

Observation 11d557d8-f6ab-4da4-96a2-e659eeccefa3 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Uncertainty-aware Reward Design Process V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.455724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.455724Z digest=sha256:0b14e752dceaffc6497c7787f8d1c45596fb8016470c225536b04978e6ce04e2

Observation e1752300-b526-4ae2-bda8-21ce563c48ec · outbound

This paper cites Empowering LLMs with Logical Reasoning: A Comprehensive Survey.

Uncertainty-aware Reward Design Process Empowering LLMs with Logical Reasoning: A Comprehensive Survey

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.583068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.583068Z digest=sha256:de83962108fa9041298853ab864975af0982be4aaf764f414740058812f8ac7f

Observation c063a9bb-689b-4bee-af8f-8e9c0d992ba0 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Uncertainty-aware Reward Design Process Language Models (Mostly) Know What They Know

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:13.821156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:13.821156Z digest=sha256:4fb3bd9e0618e7cb38c9a24c53ec602d902ac61c73af2c43b6441de08bd89afb

Observation 76f56745-1708-4541-bf02-82be042d9954 · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

Uncertainty-aware Reward Design Process Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:14.125605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:14.125605Z digest=sha256:7f24fc277e070dc70f2f9488bc3536433b9ee2e453a35e222e50b90846365580

Observation d9bca7eb-d4c8-4946-91a0-eb3102484ed8 · outbound

This paper cites an unresolved cited work.

Uncertainty-aware Reward Design Process Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:39:17.365294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:39:13.515501Z digest=sha256:db3a41f9d83383def0c229c655a205f3c3d0a21790fd572fc268ef09f769bb8a

Pith citing papers

No inbound Pith citation observations are available.