Pith. sign in

Paper Citation Record · LEDGER

Learning Explainable Dense Reward Shapes via Bayesian Optimization

As of 17 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2504.16272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16272 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.754841Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:49.831725Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:02:54.459404Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact5
  • verified fuzzy11
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be0da2d2-02ca-4c3d-a3db-62d5940f0410 · outbound

This paper cites write newline.

Learning Explainable Dense Reward Shapes via Bayesian Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.432613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.432613Z digest=sha256:e1f50b16616f35d1b74a74434f48a48f80a7bad0c6bf9599e2347844236b3d9d

Observation e2aa7bc8-7f03-4647-8dc1-d377fda7d9ef · outbound

This paper cites Searching for optimal solutions with LLM s via bayesian optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Searching for optimal solutions with LLM s via bayesian optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.864857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.440082Z digest=sha256:275e95cbb88785d0cca6003ed369e845cc87118bcae85243ec5639af7a8122a3

Observation a5a0520c-5ccb-4251-8b39-c9a7f7d2d610 · outbound

This paper cites Unexpected Improvements to Expected Improvement for Bayesian Optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unexpected Improvements to Expected Improvement for Bayesian Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.445757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.445757Z digest=sha256:47753d444d4c1cb9a9f6ef3f05236a290cfaba82e18d7af5e1bd7ee7f5ce8a99

Observation fb9fcc38-3381-4ca2-9f72-a4efe3d8e5e7 · outbound

This paper cites Bayesian optimization with llm-based acquisition functions for natural language preference elicitation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian optimization with llm-based acquisition functions for natural language preference elicitation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.451065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.451065Z digest=sha256:865ef7ec2055812fc0e389d7449da27baf62c6fd2dab68c6581b6fc2d4bd6968

Observation 2d390c19-b550-42ce-9c64-b1316cb013b9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.456487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.456487Z digest=sha256:0dfb08e4d460dad595f0fa8cd068388d21b2d827320d083ee2e8a24b5b8b4f49

Observation 15c65c35-49fa-4069-a33c-5103dcf29fc9 · outbound

This paper cites Ae: A domain-agnostic platform for adaptive experimentation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Ae: A domain-agnostic platform for adaptive experimentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.849023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.461886Z digest=sha256:76cf1e3d7374ed260f55d57cac31810bb711567297c456f2005cb3284450493b

Observation cc0f65d8-6b68-4d18-9f57-5b92b2a68c63 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Mechanistic Interpretability for AI Safety -- A Review

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.467178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.467178Z digest=sha256:b7329c264d724dc843ac269e968a136a9062b5fcd2a1afe36054960d21a9f3bb

Observation d5e0754c-e84a-4333-acba-9334c00391f3 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Enhancing reinforcement learning with dense rewards from language model critic

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.477936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.477936Z digest=sha256:9444e4c0fcd97eb2d6a595f154439c6410ad40666bf6dfb335bb2014c0a36f1f

Observation 0cf3d6ff-6d29-49ed-942a-da79774322f7 · outbound

This paper cites Dense Reward for Free in Reinforcement Learning from Human Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Dense Reward for Free in Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.482971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.482971Z digest=sha256:78013135324ce819082b148d5d889289c3654c3cf488c06cc61ae939b8fdcf8b

Observation e10fba38-328f-4081-a0ba-ee9849d7764c · outbound

This paper cites RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs.

Learning Explainable Dense Reward Shapes via Bayesian Optimization RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.489068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.489068Z digest=sha256:f0967ded5993da5c1700aadfb3e743aec1aab996b6c1a0dd5c39561a01e7c40f

Observation 26ef8640-410d-4711-8656-5ae3060cda3b · outbound

This paper cites I nstruct Z ero: Efficient instruction optimization for black-box large language models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization I nstruct Z ero: Efficient instruction optimization for black-box large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.832327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.494215Z digest=sha256:148fe5ce62f4058254d908e91f01b34e2d25c85332489e49a180bda9c642820d

Observation b500da87-2c39-4e7f-bdec-d37695cf0303 · outbound

This paper cites Improving large language models via fine-grained reinforcement learning with minimum editing constraint.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Improving large language models via fine-grained reinforcement learning with minimum editing constraint

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.499367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.499367Z digest=sha256:c2d11f169470f14e5f45984ab7490181c6f585d33175875809d691c019d54206

Observation aed75ea4-28d5-4ad9-94c2-b3cc03dd2de7 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.504385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.504385Z digest=sha256:e253d3d0e3a83dc24e468445be57d6fcbad30f4903841b44ac2189b04decec60

Observation fdae8e29-b113-42f9-8ed7-9feef875c404 · outbound

This paper cites Robust Multi-Objective Bayesian Optimization Under Input Noise.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Robust Multi-Objective Bayesian Optimization Under Input Noise

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.509845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.509845Z digest=sha256:9012460dcd6aaf2dc9153b8a012d43bed82e28b11e34f36aacef71944f128015

Observation d3ecf171-2a3b-488d-99d8-3b0ac85c3a84 · outbound

This paper cites A Comparative Study on Textual Saliency of Styles from Eye Tracking, Annotations, and Language Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization A Comparative Study on Textual Saliency of Styles from Eye Tracking, Annotations, and Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.514916Z digest=sha256:d8b9cf55f5f308744bf86ba01a6eb7adc06ba83b7b3e6c371a14570e0b97510d

Observation 92fb10e7-d4b0-4506-9d89-ab0bc274e946 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.520216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.520216Z digest=sha256:58f73212ffe35f62a229166f040dc0ddaa69ca046d1746a00a7b3c8ebc78f85c

Observation 36ab075a-59d5-49ae-8f89-c1c6a84ed3c1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.525518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.525518Z digest=sha256:1ffd543c7da60c04645c12665e6a3898700357e8367c080e41a98344a63194f5

Observation 8161ca46-facc-4825-bec4-0d5aa31e7eed · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.530577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.530577Z digest=sha256:6d9e81339ee773ac63aa89343c6772283de2f99ae1932570845d3a60450425d8

Observation d9ee4086-abdb-4475-9efa-c687835bf8f0 · outbound

This paper cites Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.351519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.535933Z digest=sha256:fb8725a56e98e0883543b47ac713cee98e4666571eeab9036a6dd38c3c761041

Observation 69800bf4-61c8-44bf-a872-39ccbd9da377 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.541053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.541053Z digest=sha256:7f33c83a17d213a7f005f0ea7cd732a5e0e6d694644d478d28b2539208facfba

Observation 21c757cb-616c-4876-9e07-fb409d191222 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Scaling Laws for Reward Model Overoptimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.546227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.546227Z digest=sha256:0874d73991e2514880288c8cb98348aa4f992e84fc6ea079d28712ed6d1997ed

Observation 071fe415-b93d-4dff-a8e6-23b4a6091094 · outbound

This paper cites B ayesian calibration of win rate estimation with LLM evaluators.

Learning Explainable Dense Reward Shapes via Bayesian Optimization B ayesian calibration of win rate estimation with LLM evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.551640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.551640Z digest=sha256:5d86ff80f3092e2c375a381392bff50ed539dd6aa0c9422da794c845d8d26049

Observation 5f8b3ec3-5902-41ba-ad03-94cc59b037b4 · outbound

This paper cites Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.556751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.556751Z digest=sha256:5076e798010d8a21989de75e2ccc7ce5c47300b3155f7a4415c6f7f553076463

Observation f9852ba3-ac7b-4ce6-af06-9581434901ab · outbound

This paper cites Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.561721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.561721Z digest=sha256:6e3239b67f17dfcf849b36fb211396182da71070a0a0885e657ff4b41f6d6af4

Observation fa33791a-c2a0-4da9-9ab5-c14e48d45fcd · outbound

This paper cites AlphaPO: Reward Shape Matters for LLM Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization AlphaPO: Reward Shape Matters for LLM Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.567250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.567250Z digest=sha256:a09f9c442c30ade51d19e3c487cb1dea1f1833532f68fc804ae64d6070b7f8f3

Observation d38521f6-a144-43f4-a909-2adc236dddbf · outbound

This paper cites Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.246578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.572878Z digest=sha256:90becc3024c9cd6a14936b5099a4be1003f823003f9451c829b4311cf00756dc

Observation ff39b7f4-85c4-408c-a901-46bbb42e1719 · outbound

This paper cites Learning to utilize shaping rewards: a new approach of reward shaping.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Learning to utilize shaping rewards: a new approach of reward shaping

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.815112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.578414Z digest=sha256:fe411a09cf69c0cd8eb6e119be2208fe0270bb187e2c4e1e96784a18c780bbaa

Observation 19a18418-60e8-44fb-a9e4-b9e688865705 · outbound

This paper cites Training language models to generate text with citations via fine-grained rewards.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Training language models to generate text with citations via fine-grained rewards

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.583131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.583131Z digest=sha256:7990d643c865b477f702fd26fb277be863d14fb8662580791d53c43e10d35ffd

Observation b16c611c-2df7-4e78-af0a-a24931d6131a · outbound

This paper cites Attention is not Explanation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Attention is not Explanation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.588294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.588294Z digest=sha256:e0ec8a69f1b4651ebfbaee97c41257ce9ca9d11f601548a26e42fadabb2411ab

Observation e51f7903-74d5-42bf-94a4-8643b79fdadc · outbound

This paper cites Align to structure: Aligning large language models with structural information, 2025.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Align to structure: Aligning large language models with structural information, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.796657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.593497Z digest=sha256:e7497ccf27292b0517d529b25e0551ee1856604ee86e6f2e9d00b029e02a2030

Observation 0ae909f1-dbc3-4b54-86f6-7cf6d7f79335 · outbound

This paper cites an unresolved cited work.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:13:54.780298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.598264Z digest=sha256:82cc75744b21663a19fbebae1fdadf7351685d67882dd136fb3ea0d7ecf345fa

Observation 4f831392-b562-4fb1-a6c7-26c59520b9f3 · outbound

This paper cites Dvornek, Yufeng Gu, Pamela Ventola, and James S.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Dvornek, Yufeng Gu, Pamela Ventola, and James S

Reference 33

Resolution
verified exact
doi, observed 2026-08-16T11:13:53.828167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.602872Z digest=sha256:ea45b4077fa56994d6b4a64254d66c1631f7255fc1b877ff0b53480116763e3d

Observation 16dc7879-2c2f-4fbc-b71b-039cffff9138 · outbound

This paper cites Let's Verify Step by Step.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Let's Verify Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.608027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.608027Z digest=sha256:ca2edd7de8fbbaf8f648e2aa3a0740c64c67919c0373d66f8d617ca47db5da11

Observation 49d47933-772f-474f-9707-fa99e5495927 · outbound

This paper cites Checkpoint merging via bayesian optimization in llm pretraining.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Checkpoint merging via bayesian optimization in llm pretraining

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.763841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.613242Z digest=sha256:3070eab38ea359efdf62e87603963c2827fd305aa3d472f06a922f54071f1de1

Observation bb2370b8-cd47-4f7c-a2be-8432c0dac0d7 · outbound

This paper cites Choosing the sample size of a computer experiment: A practical guide.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Choosing the sample size of a computer experiment: A practical guide

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.746849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.617851Z digest=sha256:5c90f9f02689e13e8b3e6efde6a9560612fdf7eaf5acd2348237fcd596d255fe

Observation 07e59041-b32c-451a-926b-d56a822a30a5 · outbound

This paper cites A unified approach to interpreting model predictions.

Learning Explainable Dense Reward Shapes via Bayesian Optimization A unified approach to interpreting model predictions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.729881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.622521Z digest=sha256:a24c434542e60510db3bfa6c0fa329b514cfa41156a13d4cd0c7aae1d2d49339

Observation 3712e313-8501-4ba4-af98-92ea12a111b0 · outbound

This paper cites López and Martha Saboyá.

Learning Explainable Dense Reward Shapes via Bayesian Optimization López and Martha Saboyá

Reference 38

Resolution
verified exact
doi, observed 2026-08-16T11:13:53.809530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.627449Z digest=sha256:39977055a105fd0099e7c1e18028f936affc8cca8a01544596c3652663a0988f

Observation ff7ced17-d3d7-4bb5-8c47-d83c83e8a7c6 · outbound

This paper cites Ng, Daishi Harada, and Stuart J.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Ng, Daishi Harada, and Stuart J

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.632376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.632376Z digest=sha256:ab90fef0f668ba48e73d748a511a584d79426ba74773115df103047daa701dcb

Observation 9d809bea-52bd-4d59-af78-2133dc328e01 · outbound

This paper cites Optimizing instructions and demonstrations for multi-stage language model programs.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Optimizing instructions and demonstrations for multi-stage language model programs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.637185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.637185Z digest=sha256:35f2f39aa4ef3c3e67a59a49dbd09be68cd9df1361f70434017cac84955c38b8

Observation d76c6724-72a8-40b7-bd7c-a561ccbe78c7 · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level Proximal Policy Optimization for Query Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.641848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.641848Z digest=sha256:19868626f778d90e72a4aab117d9557617b5b0fd8c9d987aaba687612e7cbb9c

Observation bf6547e5-2a7c-42c0-a919-fb03946209ea · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Learning Explainable Dense Reward Shapes via Bayesian Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.647127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.647127Z digest=sha256:c57aae6b6dfa3a6d3440ed5770f1f28b1bb8a76bbd81bd40de6423d3914ef23d

Observation b0101030-01d9-421d-8a35-a662453e089f · outbound

This paper cites Vanishing Gradients in Reinforcement Finetuning of Language Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Vanishing Gradients in Reinforcement Finetuning of Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.652373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.652373Z digest=sha256:0f1190711dd540a0a78ef998cda44cffda1433655ed3e5146e61f3852c2ff888

Observation 8d3a9fe6-74ed-4604-8fd6-8bb6c2d72a41 · outbound

This paper cites "Why Should I Trust You?": Explaining the Predictions of Any Classifier.

Learning Explainable Dense Reward Shapes via Bayesian Optimization "Why Should I Trust You?": Explaining the Predictions of Any Classifier

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.657392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.657392Z digest=sha256:20f654157c4f8d40f24f96ef22e2b0744d82d3b69c3dd25f4e82e5ae9f3c4a6e

Observation 6640c93f-2c79-4efc-b5be-ef0233c98c4f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Proximal Policy Optimization Algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.662760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.662760Z digest=sha256:cd9be9d1fa1fcb8a910c4aece86395ee561f3e509cb45a1e80cb2297aa123aae

Observation 893e5ac8-0439-4968-98f4-9431b38b94b5 · outbound

This paper cites Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.668079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.668079Z digest=sha256:b6765803227b512c07313b0f4e7742bb198ef349b2a2a3195a0169bda83c7bab

Observation 240ed774-ab87-43bb-b7e8-f10721d1bd13 · outbound

This paper cites Practical Bayesian Optimization of Machine Learning Algorithms.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Practical Bayesian Optimization of Machine Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.672974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.672974Z digest=sha256:6c1fb6d36cd33c0a95ae84a88a06ee00ebe9c2bc51a3dc4b197365d6c1f67f4b

Observation 8185d6af-748d-456a-bfd7-838dc0ef45a8 · outbound

This paper cites Sutton and Andrew G.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Sutton and Andrew G

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.678405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.678405Z digest=sha256:41dd94d63b0abbca9a5b6effb159e42875cbdd4e7296dfbdba4b4a2a07cadba3

Observation 924813ae-32c8-4ac8-9beb-546af64e8672 · outbound

This paper cites The Llama 3 Herd of Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization The Llama 3 Herd of Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.683188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.683188Z digest=sha256:3e09894b15d13ea73b240c53d2b509eec1cb5b4df713570a80920ee95bcca856

Observation 09b66e02-8d0e-4e46-9e1e-9d0b9ce44ba2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Solving math word problems with process- and outcome-based feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.688401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.688401Z digest=sha256:a233c7d22e8bf68fb39bdc5e1d439ee04b0f6fb099dc8e1b8bd13f02e8db7399

Observation 14fca40d-d3cd-49ad-a4c6-cfb0410a079b · outbound

This paper cites Trl: Transformer reinforcement learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Trl: Transformer reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.693486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.693486Z digest=sha256:84546d070d7e176daad0f86064f8481b8802033ecbd21d396c4fb964564d7b8d

Observation b8be739a-47c1-491d-9a4b-64c8c36d1227 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Fine-Grained Human Feedback Gives Better Rewards for Language Model Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.698400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.698400Z digest=sha256:3fc39bd079aa070e2cfe491745e77ef59bb7fec028799db4b0c98e01dd8f7d6d

Observation 1757d1cf-61d8-4ada-9e0c-f1774b93a150 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.703608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.703608Z digest=sha256:5e946a864a8e7f5b344224ae48794ac0c47b97d72e1334995314cbf6f477d515

Observation f387338b-f995-4df2-8be2-8cc04e0e53b2 · outbound

This paper cites Bayesian reward models for llm alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian reward models for llm alignment

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.681298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.708866Z digest=sha256:7d5f63a59ebcdbbbf1e6a178fc122b37baf46819fa91d6a92819ada121b76035

Observation da0bccb0-6288-4891-90e1-995df1dfcd19 · outbound

This paper cites Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.664323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.713467Z digest=sha256:4d241a2cbdf98254ad3c8e216593bb60c709f71fe9c0464e80cbedcb93c2e180

Observation 4c7b430a-b5cd-4705-86c5-900022d99dfc · outbound

This paper cites Token-level direct preference optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level direct preference optimization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.647415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.718102Z digest=sha256:752687cb7581b3016dbabe12d65a7d2990b5348f6bcfe75c740ab58b68eba60e

Observation 343a2e57-0c8c-4d85-a8d1-03d312a63d27 · outbound

This paper cites An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.722905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.722905Z digest=sha256:e7cfcada4f44489bfaad035b12746a14f9aa404fa1d8309f00322ec4860363b2

Observation dcab4207-4662-4744-8371-55243852e5c2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.728896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.728896Z digest=sha256:617366a4e34b3e793a73310537b021f931a98d93093f8d5c41197809d7162f09

Observation aa4cbb0e-1b32-42ba-91b7-8c444343372b · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Secrets of RLHF in Large Language Models Part I: PPO

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.734394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.734394Z digest=sha256:02eb97979d5225d90755c7b879699e1be1350eedf021869bf8eb830bb8d624bc

Observation 1dc7d0db-9095-4dc4-bb4b-a1b3ed8c2c08 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.739536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.739536Z digest=sha256:4c29017a90e400505b7fd67ef1d1971fea95667ad690df532dd9d428b0c1fe21

Observation 6d3e2307-1506-41eb-ae9e-126c16d9f2a9 · outbound

This paper cites @esa (Ref.

Learning Explainable Dense Reward Shapes via Bayesian Optimization @esa (Ref

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.744755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.744755Z digest=sha256:f184a9f18e65380a52e5a30e2d5d89e1e6a856a7c73e658b2a8987d1d8b3b732

Observation 4531159c-9783-4a26-a212-d2345a7584ef · outbound

This paper cites an unresolved cited work.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.749803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.749803Z digest=sha256:1af1452008c85cc581f67f8bc026fe01532dc999db75e3326b62bfe2eb790d3f

Observation a1fc9a04-f739-4ead-bd27-43f969e7b283 · outbound

This paper cites reward shaping.

Learning Explainable Dense Reward Shapes via Bayesian Optimization reward shaping

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.754841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.754841Z digest=sha256:c1d6f59e10827e997406ae46f5c109794ee24c3638335af114c47f1210c09102

Pith citing papers

Observation 923e16fe-d3d4-41ca-817b-2a72a70795c5 · inbound

SCAR: Shapley Credit Assignment for More Efficient RLHF cites this paper.

SCAR: Shapley Credit Assignment for More Efficient RLHF Learning Explainable Dense Reward Shapes via Bayesian Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:02:54.542906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:49.831725Z digest=sha256:b9f17aa3fdc89af33f44592f3dbd2c670730a4b59433efb7072b9cff2176c433