Pith. sign in

Paper Citation Record · LEDGER

Learning Explainable Dense Reward Shapes via Bayesian Optimization

As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2504.16272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16272 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.754841Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:49.831725Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:02:54.459404Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact5
  • verified fuzzy11
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be0da2d2-02ca-4c3d-a3db-62d5940f0410 · outbound

This paper cites write newline.

Learning Explainable Dense Reward Shapes via Bayesian Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.432613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.432613Z digest=sha256:629aeea59cafc53e335942943fbcbab96783debc8f4716afbc62133e9a99797a

Observation e2aa7bc8-7f03-4647-8dc1-d377fda7d9ef · outbound

This paper cites Searching for optimal solutions with LLM s via bayesian optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Searching for optimal solutions with LLM s via bayesian optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.864857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.440082Z digest=sha256:0ee25bffc00cf46d833a03c9b292e2ddab1a0871c9794500f10df6345efdd871

Observation a5a0520c-5ccb-4251-8b39-c9a7f7d2d610 · outbound

This paper cites Unexpected Improvements to Expected Improvement for Bayesian Optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unexpected Improvements to Expected Improvement for Bayesian Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.445757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.445757Z digest=sha256:94b06857eb804d056a3e22df370ed5bfecfae8ce57a2686a94c6f6221776c101

Observation fb9fcc38-3381-4ca2-9f72-a4efe3d8e5e7 · outbound

This paper cites Bayesian optimization with llm-based acquisition functions for natural language preference elicitation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian optimization with llm-based acquisition functions for natural language preference elicitation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.451065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.451065Z digest=sha256:a69b65d59e86294bb998dccdf422712b5505f8666fa2d3a094d0ef0a360df340

Observation 2d390c19-b550-42ce-9c64-b1316cb013b9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.456487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.456487Z digest=sha256:d2edcea299007bf71a0810282bbb1078bd58126c9d0fe95a2dc07bdf921cc173

Observation 15c65c35-49fa-4069-a33c-5103dcf29fc9 · outbound

This paper cites Ae: A domain-agnostic platform for adaptive experimentation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Ae: A domain-agnostic platform for adaptive experimentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.849023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.461886Z digest=sha256:97c57354ef195e82a65923703f19331b7f5930b0775b3ed03af9bc01ead4ec30

Observation cc0f65d8-6b68-4d18-9f57-5b92b2a68c63 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Mechanistic Interpretability for AI Safety -- A Review

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.467178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.467178Z digest=sha256:b426ce4d38dbcfe440f5da39bdd8f543c2a6c4c304b439d54087175e544bf38d

Observation d5e0754c-e84a-4333-acba-9334c00391f3 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Enhancing reinforcement learning with dense rewards from language model critic

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.477936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.477936Z digest=sha256:1296c9fe90cefee17de69ceea387de21a186bfc1325eca74d02c9ad585e2ca3b

Observation 0cf3d6ff-6d29-49ed-942a-da79774322f7 · outbound

This paper cites Dense Reward for Free in Reinforcement Learning from Human Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Dense Reward for Free in Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.482971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.482971Z digest=sha256:cf5866d3a449ceb8ff63433b3f92599529342cf6f59d2e233541ddfc6485e0d4

Observation e10fba38-328f-4081-a0ba-ee9849d7764c · outbound

This paper cites RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs.

Learning Explainable Dense Reward Shapes via Bayesian Optimization RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.489068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.489068Z digest=sha256:406bb4e7709cca24b7b720bb1c84be9523d70cff14a109f861db8a312b30b63c

Observation 26ef8640-410d-4711-8656-5ae3060cda3b · outbound

This paper cites I nstruct Z ero: Efficient instruction optimization for black-box large language models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization I nstruct Z ero: Efficient instruction optimization for black-box large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.832327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.494215Z digest=sha256:39668782de459e8c007b2847eece8ca5ec43eab63479afd3ad4185b27e2bd599

Observation b500da87-2c39-4e7f-bdec-d37695cf0303 · outbound

This paper cites Improving large language models via fine-grained reinforcement learning with minimum editing constraint.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Improving large language models via fine-grained reinforcement learning with minimum editing constraint

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.499367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.499367Z digest=sha256:4dac00568a16258c4b2b444de2c6f748ee7461f7c80a78db90b02cbe65c66041

Observation aed75ea4-28d5-4ad9-94c2-b3cc03dd2de7 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.504385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.504385Z digest=sha256:cc0aad5a6817b35748912c3a9808d908ba0c4e66a17db85362767886e93f7af0

Observation fdae8e29-b113-42f9-8ed7-9feef875c404 · outbound

This paper cites Robust Multi-Objective Bayesian Optimization Under Input Noise.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Robust Multi-Objective Bayesian Optimization Under Input Noise

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.509845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.509845Z digest=sha256:162fbd44b22ae031e3fa98950fa3bd3f31966425488cf618e160bf62ba0eadc8

Observation d3ecf171-2a3b-488d-99d8-3b0ac85c3a84 · outbound

This paper cites A Comparative Study on Textual Saliency of Styles from Eye Tracking, Annotations, and Language Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization A Comparative Study on Textual Saliency of Styles from Eye Tracking, Annotations, and Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.514916Z digest=sha256:0266724c4f5a3de51394a1eb2e87c0af7ccede8a7241861dab91775f591b2c4e

Observation 92fb10e7-d4b0-4506-9d89-ab0bc274e946 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.520216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.520216Z digest=sha256:a21027cd5bc513fe553b63b6609cb0092c6eec56d258f9e91803dddc64e041f1

Observation 36ab075a-59d5-49ae-8f89-c1c6a84ed3c1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.525518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.525518Z digest=sha256:77c3cd9f88cf10df217d1c2b57cc39e8cfc8ac6d801e8f54420b34d0e8c6be0a

Observation 8161ca46-facc-4825-bec4-0d5aa31e7eed · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.530577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.530577Z digest=sha256:dd1d5b14c8556b4c6a476899096fefc615555138edf8fbe3987ea28bb4a23b4c

Observation d9ee4086-abdb-4475-9efa-c687835bf8f0 · outbound

This paper cites Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.351519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.535933Z digest=sha256:7d48ffacb43a759c208370d29e084bd3d65d0d3e444cd5447cdfe3fc8f272f9e

Observation 69800bf4-61c8-44bf-a872-39ccbd9da377 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.541053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.541053Z digest=sha256:bd20bdcd4b4c3e2e606b9e22ddb304cce2a213bb62f94b0eba45232de75abdf5

Observation 21c757cb-616c-4876-9e07-fb409d191222 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Scaling Laws for Reward Model Overoptimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.546227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.546227Z digest=sha256:38e4b700568ff7f3141175b9d6ade37d750e85ab3347ab1bdb34d6adc9953a88

Observation 071fe415-b93d-4dff-a8e6-23b4a6091094 · outbound

This paper cites B ayesian calibration of win rate estimation with LLM evaluators.

Learning Explainable Dense Reward Shapes via Bayesian Optimization B ayesian calibration of win rate estimation with LLM evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.551640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.551640Z digest=sha256:398d88a2a4d5e39ee545fe0fd6ec3289c605f5f34e5f132bdb70d7c409336337

Observation 5f8b3ec3-5902-41ba-ad03-94cc59b037b4 · outbound

This paper cites Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.556751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.556751Z digest=sha256:b86125b7216c08ae81f0799d1c63b9ceb362161cbc90da74fe25f082eb6b55c5

Observation f9852ba3-ac7b-4ce6-af06-9581434901ab · outbound

This paper cites Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.561721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.561721Z digest=sha256:1504edda2712b07528dc5e2da295bb2d749980c111f0aad64e52818530e5f53e

Observation fa33791a-c2a0-4da9-9ab5-c14e48d45fcd · outbound

This paper cites AlphaPO: Reward Shape Matters for LLM Alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization AlphaPO: Reward Shape Matters for LLM Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.567250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.567250Z digest=sha256:abda7bd7e390b5cf8cd1e3a50c0fc0c00aa156cd3165e764fc2d4264aa17a685

Observation d38521f6-a144-43f4-a909-2adc236dddbf · outbound

This paper cites Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:13:54.246578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.572878Z digest=sha256:686677f9535af6046dd8a93b9a3e82be658b6479b947ad21005f8d02624d8256

Observation ff39b7f4-85c4-408c-a901-46bbb42e1719 · outbound

This paper cites Learning to utilize shaping rewards: a new approach of reward shaping.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Learning to utilize shaping rewards: a new approach of reward shaping

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.815112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.578414Z digest=sha256:285a2aef722583b64bb5470a9ab0bd30f98d147b343948623ec995f07ba3460a

Observation 19a18418-60e8-44fb-a9e4-b9e688865705 · outbound

This paper cites Training language models to generate text with citations via fine-grained rewards.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Training language models to generate text with citations via fine-grained rewards

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.583131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.583131Z digest=sha256:fb0e086923d79154be32cd5481eda8aabd3195269702b1ff20148ce8dceb4980

Observation b16c611c-2df7-4e78-af0a-a24931d6131a · outbound

This paper cites Attention is not Explanation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Attention is not Explanation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.588294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.588294Z digest=sha256:08de3f2cf27bce72f9eeafe437a014ed6ed56c0ddfb7b43bb6db59fc5e12191c

Observation e51f7903-74d5-42bf-94a4-8643b79fdadc · outbound

This paper cites Align to structure: Aligning large language models with structural information, 2025.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Align to structure: Aligning large language models with structural information, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.796657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.593497Z digest=sha256:86b2d63b376a063cae9427d8df45f99b7bed77308789876675ed4ac488dff06f

Observation 0ae909f1-dbc3-4b54-86f6-7cf6d7f79335 · outbound

This paper cites an unresolved cited work.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:13:54.780298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.598264Z digest=sha256:6e39ddd19f13b2e09e0ae6026d0fe6972022cebc8de32c0899c5cd43c3c206c4

Observation 4f831392-b562-4fb1-a6c7-26c59520b9f3 · outbound

This paper cites Dvornek, Yufeng Gu, Pamela Ventola, and James S.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Dvornek, Yufeng Gu, Pamela Ventola, and James S

Reference 33

Resolution
verified exact
doi, observed 2026-08-16T11:13:53.828167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.602872Z digest=sha256:d4b30a42bf799dde0156048b9d7e4e6b10f92d5f98f5d0eb4d0398d9941591ee

Observation 16dc7879-2c2f-4fbc-b71b-039cffff9138 · outbound

This paper cites Let's Verify Step by Step.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Let's Verify Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.608027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.608027Z digest=sha256:0f22ff4da0cda5c70c82842711fe466c950e08b816656223265055177b15c281

Observation 49d47933-772f-474f-9707-fa99e5495927 · outbound

This paper cites Checkpoint merging via bayesian optimization in llm pretraining.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Checkpoint merging via bayesian optimization in llm pretraining

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.763841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.613242Z digest=sha256:67ada1755a6b8de4fdc858a20fa1c79e999f3da6efbf26193e5a08bc3d837201

Observation bb2370b8-cd47-4f7c-a2be-8432c0dac0d7 · outbound

This paper cites Choosing the sample size of a computer experiment: A practical guide.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Choosing the sample size of a computer experiment: A practical guide

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.746849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.617851Z digest=sha256:2f27e8f60555b1622b46d1814d8ac97dbb7b49092ba0bd42ea1c13bc9da6a978

Observation 07e59041-b32c-451a-926b-d56a822a30a5 · outbound

This paper cites A unified approach to interpreting model predictions.

Learning Explainable Dense Reward Shapes via Bayesian Optimization A unified approach to interpreting model predictions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.729881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.622521Z digest=sha256:c98e758613f7173abfe5afda122a665480a941f13f0e2142d95220630a8f749c

Observation 3712e313-8501-4ba4-af98-92ea12a111b0 · outbound

This paper cites López and Martha Saboyá.

Learning Explainable Dense Reward Shapes via Bayesian Optimization López and Martha Saboyá

Reference 38

Resolution
verified exact
doi, observed 2026-08-16T11:13:53.809530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.627449Z digest=sha256:0539490ed13e12d8a06a6ce942c5612038c6029505eddb371afdd4187c51c796

Observation ff7ced17-d3d7-4bb5-8c47-d83c83e8a7c6 · outbound

This paper cites Ng, Daishi Harada, and Stuart J.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Ng, Daishi Harada, and Stuart J

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.632376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.632376Z digest=sha256:aa05d9b137cdb3d79b908e8b65f9d11d787f7a065fecbc1c54228e6cd0546fd4

Observation 9d809bea-52bd-4d59-af78-2133dc328e01 · outbound

This paper cites Optimizing instructions and demonstrations for multi-stage language model programs.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Optimizing instructions and demonstrations for multi-stage language model programs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.637185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.637185Z digest=sha256:e4b1f326901ebe52719b5f8a290dd31111209a125a3fc789cd6e4d82aaf84e14

Observation d76c6724-72a8-40b7-bd7c-a561ccbe78c7 · outbound

This paper cites Token-level Proximal Policy Optimization for Query Generation.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level Proximal Policy Optimization for Query Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.641848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.641848Z digest=sha256:9ba330d733543358810cad1fa8cdb75608f4835b5b99f301260468ecebc5cc3a

Observation bf6547e5-2a7c-42c0-a919-fb03946209ea · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Learning Explainable Dense Reward Shapes via Bayesian Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.647127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.647127Z digest=sha256:68ac320f6a9784470bddb1cfce143617383cbef1b35657eea0342d19117b4e4b

Observation b0101030-01d9-421d-8a35-a662453e089f · outbound

This paper cites Vanishing Gradients in Reinforcement Finetuning of Language Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Vanishing Gradients in Reinforcement Finetuning of Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.652373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.652373Z digest=sha256:4514537280914177274d454c12eb93d52d489258dcced1c6e0e88138512b868b

Observation 8d3a9fe6-74ed-4604-8fd6-8bb6c2d72a41 · outbound

This paper cites "Why Should I Trust You?": Explaining the Predictions of Any Classifier.

Learning Explainable Dense Reward Shapes via Bayesian Optimization "Why Should I Trust You?": Explaining the Predictions of Any Classifier

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.657392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.657392Z digest=sha256:1f773266124265084151a7ad6bdc974b179142a5b1bc956795949a91efb50e15

Observation 6640c93f-2c79-4efc-b5be-ef0233c98c4f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Proximal Policy Optimization Algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.662760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.662760Z digest=sha256:7636ac693e2d78655ad0346674f5db45abcfa31955b3e0d026da6ee638241599

Observation 893e5ac8-0439-4968-98f4-9431b38b94b5 · outbound

This paper cites Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.668079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.668079Z digest=sha256:e933ef85d29fecc410d4ba4059e4e0a467d0b36c546e7125c9039edce6398e8e

Observation 240ed774-ab87-43bb-b7e8-f10721d1bd13 · outbound

This paper cites Practical Bayesian Optimization of Machine Learning Algorithms.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Practical Bayesian Optimization of Machine Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.672974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.672974Z digest=sha256:6da68a0b275543ba728f18e95864024fd01439fe04e90af8c265a9d5d9d04f26

Observation 8185d6af-748d-456a-bfd7-838dc0ef45a8 · outbound

This paper cites Sutton and Andrew G.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Sutton and Andrew G

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.678405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.678405Z digest=sha256:da459c76d9f0a8e66dadd9a4d37450eec3049d3e673bef775b72a0469d0388fc

Observation 924813ae-32c8-4ac8-9beb-546af64e8672 · outbound

This paper cites The Llama 3 Herd of Models.

Learning Explainable Dense Reward Shapes via Bayesian Optimization The Llama 3 Herd of Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.683188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.683188Z digest=sha256:4467aca8b9defaf7e23e7135f99a72a85917eceaaecb8476e0960fe062e8a65b

Observation 09b66e02-8d0e-4e46-9e1e-9d0b9ce44ba2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Solving math word problems with process- and outcome-based feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.688401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.688401Z digest=sha256:59723e91a260c57c5681fcc6acf8790f13eae5d535a59f14c7a554aa24b2dea1

Observation 14fca40d-d3cd-49ad-a4c6-cfb0410a079b · outbound

This paper cites Trl: Transformer reinforcement learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Trl: Transformer reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.693486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.693486Z digest=sha256:5140b72c34bf9ab3ab3c1cdd31da2de787553f77fd1ef17181803b58fe152511

Observation b8be739a-47c1-491d-9a4b-64c8c36d1227 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Fine-Grained Human Feedback Gives Better Rewards for Language Model Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.698400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.698400Z digest=sha256:d9e0d680df3c1e22f5f60e4b7ac197bd236bd99732357e738699799710310ba3

Observation 1757d1cf-61d8-4ada-9e0c-f1774b93a150 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.703608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.703608Z digest=sha256:13c7f11e20d74fe8a5749dfddddcc82e4f92587f932c8e6cb28bdf1823c804b1

Observation f387338b-f995-4df2-8be2-8cc04e0e53b2 · outbound

This paper cites Bayesian reward models for llm alignment.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian reward models for llm alignment

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.681298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.708866Z digest=sha256:96820ad4b1b1e570d581e770560d3e092c22caeaeacfcd0c27707c53f5a2f02d

Observation da0bccb0-6288-4891-90e1-995df1dfcd19 · outbound

This paper cites Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.664323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.713467Z digest=sha256:caa4b66bc195b462700203f65bdb88cc33e709ff5fabd56c094f5ff0f7add1e7

Observation 4c7b430a-b5cd-4705-86c5-900022d99dfc · outbound

This paper cites Token-level direct preference optimization.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level direct preference optimization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:13:54.647415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:13:53.718102Z digest=sha256:10bb68157954ad772cf8d1dc65df88cc7d341b2b3814e235bf4d9a319e78157e

Observation 343a2e57-0c8c-4d85-a8d1-03d312a63d27 · outbound

This paper cites An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning.

Learning Explainable Dense Reward Shapes via Bayesian Optimization An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.722905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.722905Z digest=sha256:b99e2e7e3582deddea46f8320142be3e4b2adcfe7cf2dfb4f638332e948a9cdc

Observation dcab4207-4662-4744-8371-55243852e5c2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.728896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.728896Z digest=sha256:37833980f4627ed9dc753bf34c7a79d9237c7f7dc069b06ac5267ac566f54c4f

Observation aa4cbb0e-1b32-42ba-91b7-8c444343372b · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Secrets of RLHF in Large Language Models Part I: PPO

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.734394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.734394Z digest=sha256:698f5055ddcc5fb9f26c103a648337c137c6b702d3a02384961975ea66ba5fde

Observation 1dc7d0db-9095-4dc4-bb4b-a1b3ed8c2c08 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Learning Explainable Dense Reward Shapes via Bayesian Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.739536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.739536Z digest=sha256:831c16e14ec215268571c7b4ee0fe03ae8dd037c3e3e95fb806ca618e3e86f6d

Observation 6d3e2307-1506-41eb-ae9e-126c16d9f2a9 · outbound

This paper cites @esa (Ref.

Learning Explainable Dense Reward Shapes via Bayesian Optimization @esa (Ref

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.744755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.744755Z digest=sha256:52cae8d4149bc93e658e7f020a788e544403ed82ec41d34aa95ae7f6c4e0d6c9

Observation 4531159c-9783-4a26-a212-d2345a7584ef · outbound

This paper cites an unresolved cited work.

Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.749803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.749803Z digest=sha256:603707ebb5d7d55ca0a72157908a9b9ec30e72656c295ed11a4f24b450b98c00

Observation a1fc9a04-f739-4ead-bd27-43f969e7b283 · outbound

This paper cites reward shaping.

Learning Explainable Dense Reward Shapes via Bayesian Optimization reward shaping

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.754841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.754841Z digest=sha256:f5536a8dd4f7c51ea08c657389e5bade5cdde0004c7941f34875caca27e3a3d0

Pith citing papers

Observation 923e16fe-d3d4-41ca-817b-2a72a70795c5 · inbound

SCAR: Shapley Credit Assignment for More Efficient RLHF cites this paper.

SCAR: Shapley Credit Assignment for More Efficient RLHF Learning Explainable Dense Reward Shapes via Bayesian Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:02:54.542906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:02:49.831725Z digest=sha256:4f797e99d5540741b9253ea745c9bc93bc218c3c0a0177e0023a4d6901c2839a