Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.754841Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2504.16272.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.754841Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:49.831725Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:02:54.459404Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be0da2d2-02ca-4c3d-a3db-62d5940f0410 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2aa7bc8-7f03-4647-8dc1-d377fda7d9ef · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Searching for optimal solutions with LLM s via bayesian optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a5a0520c-5ccb-4251-8b39-c9a7f7d2d610 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Unexpected Improvements to Expected Improvement for Bayesian Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9fcc38-3381-4ca2-9f72-a4efe3d8e5e7 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian optimization with llm-based acquisition functions for natural language preference elicitation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d390c19-b550-42ce-9c64-b1316cb013b9 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c65c35-49fa-4069-a33c-5103dcf29fc9 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Ae: A domain-agnostic platform for adaptive experimentation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc0f65d8-6b68-4d18-9f57-5b92b2a68c63 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Mechanistic Interpretability for AI Safety -- A Review
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e0754c-e84a-4333-acba-9334c00391f3 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Enhancing reinforcement learning with dense rewards from language model critic
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf3d6ff-6d29-49ed-942a-da79774322f7 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10fba38-328f-4081-a0ba-ee9849d7764c · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ef8640-410d-4711-8656-5ae3060cda3b · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization I nstruct Z ero: Efficient instruction optimization for black-box large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b500da87-2c39-4e7f-bdec-d37695cf0303 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Improving large language models via fine-grained reinforcement learning with minimum editing constraint
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed75ea4-28d5-4ad9-94c2-b3cc03dd2de7 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdae8e29-b113-42f9-8ed7-9feef875c404 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Robust Multi-Objective Bayesian Optimization Under Input Noise
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ecf171-2a3b-488d-99d8-3b0ac85c3a84 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization A Comparative Study on Textual Saliency of Styles from Eye Tracking, Annotations, and Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92fb10e7-d4b0-4506-9d89-ab0bc274e946 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ab075a-59d5-49ae-8f89-c1c6a84ed3c1 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8161ca46-facc-4825-bec4-0d5aa31e7eed · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ee4086-abdb-4475-9efa-c687835bf8f0 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 69800bf4-61c8-44bf-a872-39ccbd9da377 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c757cb-616c-4876-9e07-fb409d191222 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Scaling Laws for Reward Model Overoptimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 071fe415-b93d-4dff-a8e6-23b4a6091094 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization B ayesian calibration of win rate estimation with LLM evaluators
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8b3ec3-5902-41ba-ad03-94cc59b037b4 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9852ba3-ac7b-4ce6-af06-9581434901ab · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa33791a-c2a0-4da9-9ab5-c14e48d45fcd · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization AlphaPO: Reward Shape Matters for LLM Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38521f6-a144-43f4-a909-2adc236dddbf · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff39b7f4-85c4-408c-a901-46bbb42e1719 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Learning to utilize shaping rewards: a new approach of reward shaping
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19a18418-60e8-44fb-a9e4-b9e688865705 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Training language models to generate text with citations via fine-grained rewards
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16c611c-2df7-4e78-af0a-a24931d6131a · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Attention is not Explanation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51f7903-74d5-42bf-94a4-8643b79fdadc · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Align to structure: Aligning large language models with structural information, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ae909f1-dbc3-4b54-86f6-7cf6d7f79335 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f831392-b562-4fb1-a6c7-26c59520b9f3 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Dvornek, Yufeng Gu, Pamela Ventola, and James S
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 16dc7879-2c2f-4fbc-b71b-039cffff9138 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Let's Verify Step by Step
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d47933-772f-474f-9707-fa99e5495927 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Checkpoint merging via bayesian optimization in llm pretraining
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bb2370b8-cd47-4f7c-a2be-8432c0dac0d7 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Choosing the sample size of a computer experiment: A practical guide
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07e59041-b32c-451a-926b-d56a822a30a5 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization A unified approach to interpreting model predictions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3712e313-8501-4ba4-af98-92ea12a111b0 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization López and Martha Saboyá
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff7ced17-d3d7-4bb5-8c47-d83c83e8a7c6 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Ng, Daishi Harada, and Stuart J
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d809bea-52bd-4d59-af78-2133dc328e01 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Optimizing instructions and demonstrations for multi-stage language model programs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76c6724-72a8-40b7-bd7c-a561ccbe78c7 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level Proximal Policy Optimization for Query Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6547e5-2a7c-42c0-a919-fb03946209ea · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0101030-01d9-421d-8a35-a662453e089f · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Vanishing Gradients in Reinforcement Finetuning of Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3a9fe6-74ed-4604-8fd6-8bb6c2d72a41 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization "Why Should I Trust You?": Explaining the Predictions of Any Classifier
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6640c93f-2c79-4efc-b5be-ef0233c98c4f · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Proximal Policy Optimization Algorithms
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 893e5ac8-0439-4968-98f4-9431b38b94b5 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240ed774-ab87-43bb-b7e8-f10721d1bd13 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Practical Bayesian Optimization of Machine Learning Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8185d6af-748d-456a-bfd7-838dc0ef45a8 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Sutton and Andrew G
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924813ae-32c8-4ac8-9beb-546af64e8672 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization The Llama 3 Herd of Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b66e02-8d0e-4e46-9e1e-9d0b9ce44ba2 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Solving math word problems with process- and outcome-based feedback
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14fca40d-d3cd-49ad-a4c6-cfb0410a079b · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Trl: Transformer reinforcement learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8be739a-47c1-491d-9a4b-64c8c36d1227 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1757d1cf-61d8-4ada-9e0c-f1774b93a150 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f387338b-f995-4df2-8be2-8cc04e0e53b2 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Bayesian reward models for llm alignment
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da0bccb0-6288-4891-90e1-995df1dfcd19 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c7b430a-b5cd-4705-86c5-900022d99dfc · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Token-level direct preference optimization
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 343a2e57-0c8c-4d85-a8d1-03d312a63d27 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcab4207-4662-4744-8371-55243852e5c2 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4cbb0e-1b32-42ba-91b7-8c444343372b · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Secrets of RLHF in Large Language Models Part I: PPO
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc7d0db-9095-4dc4-bb4b-a1b3ed8c2c08 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3e2307-1506-41eb-ae9e-126c16d9f2a9 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization @esa (Ref
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4531159c-9783-4a26-a212-d2345a7584ef · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1fc9a04-f739-4ead-bd27-43f969e7b283 · outbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization reward shaping
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923e16fe-d3d4-41ca-817b-2a72a70795c5 · inbound
SCAR: Shapley Credit Assignment for More Efficient RLHF Learning Explainable Dense Reward Shapes via Bayesian Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.