Pith. sign in

Paper Citation Record · LEDGER

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2507.05913.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05913 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:28:59.743545Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:27:28.637485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T20:39:29.243714Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ce6e034-1836-4f15-b008-c71c9bb35f0e · outbound

This paper cites write newline.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:52.559956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:52.559956Z digest=sha256:831119f4a60d4c88867d69aab6254c45890d687dc46593570f3eaf30f4328237

Observation d04662fd-b050-4620-bb54-e25119632517 · outbound

This paper cites GPT-4 Technical Report.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:52.618979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:52.618979Z digest=sha256:b6ae52a18fbf15dfad24f12d7be326c818bbdb8f5a0d3f2ebfb3c355d1e1fded

Observation 3ebd86b7-d44a-4208-906b-59d51d5aecf8 · outbound

This paper cites Variational best-of-n alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Variational best-of-n alignment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:07.511226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:52.705364Z digest=sha256:00158749ed7bb797409a208e7579c3c95af5747293ad15b8deb64ed5226c68f4

Observation 601540f5-577a-44bd-969b-b44d2261e935 · outbound

This paper cites Theoretical analysis of kl-regularized rlhf with multiple reference models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Theoretical analysis of kl-regularized rlhf with multiple reference models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:52.821009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:52.821009Z digest=sha256:874866ac91f25122e0d2e8a22e1ede3777bc5c3b2b659587e1b7c610e4ebfc37

Observation b1af68a8-0e1c-43bf-a1fd-43ded7d1787d · outbound

This paper cites Concrete Problems in AI Safety.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Concrete Problems in AI Safety

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:52.913576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:52.913576Z digest=sha256:3e39b906b04617fe4bb3e95a931fa31cd049b302209edecaff10106f3e1c6938

Observation 19bfc809-1ae5-4de0-bbfc-a668ae52244c · outbound

This paper cites Infalign: Inference-aware language model alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Infalign: Inference-aware language model alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:07.243477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:53.001945Z digest=sha256:f3b6f199e5c055b46ec371c615bba1c36c20334b805c600a0a93a7e07ad20a42

Observation 2472b580-9c99-416e-b729-29f25acd089f · outbound

This paper cites Theoretical guarantees on the best-of-n alignment policy.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Theoretical guarantees on the best-of-n alignment policy

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:07.019005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:53.122617Z digest=sha256:06a6e908820419c4089868ba4aada5dd41958b9fa42563e8c436d26faeea14df

Observation 161a9e20-845a-48ee-839d-021f3162647f · outbound

This paper cites Q-learning for risk-sensitive control.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Q-learning for risk-sensitive control

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:06.732725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:53.203873Z digest=sha256:00d98be39d169ff954bce352385987eb161d7c173d9737ddaa6a3157cb742b40

Observation 88478503-8ef4-4b14-9a3e-694e37266884 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:53.329718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:53.329718Z digest=sha256:0e2cc0fb79bc6f846d759a01c3de5797ea126cd2d6108f382375098fcfd9ba60

Observation 3bad5a91-22c5-419d-8c94-a3902c77982d · outbound

This paper cites A short note on an inequality between KL and TV.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A short note on an inequality between KL and TV

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:53.449165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:53.449165Z digest=sha256:7cad1d3afd0e1157faa5ab7c931ace49be17f18ca53dcbf04eb9e243951b9668

Observation 5aa90d68-8dd2-4b55-8f6e-92ca8bc28ee2 · outbound

This paper cites The master equation and the convergence problem in mean field games:(ams-201).

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis The master equation and the convergence problem in mean field games:(ams-201)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:06.473880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:53.524864Z digest=sha256:546dc26b2d47b82d74dcf9435952b34b702e0816237aac8b270c680c761ea0bf

Observation ac2747d2-67f7-42a0-9154-a4e87dbab96e · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:53.661417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:53.661417Z digest=sha256:38ca7aaa21d9af4f3b00b356e0314272bd8c002a0d9bf54c5769290690fb4027

Observation bb73da34-49fc-4bb7-908b-50f43ba162e5 · outbound

This paper cites Deep reinforcement learning from human preferences.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:53.798241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:53.798241Z digest=sha256:994e535eb7faf86c04fefb72441066be1a03e60b2db4283b6e4aff19ce0b901e

Observation daf9ad92-3588-4f28-b6ce-f0ab9bda33f9 · outbound

This paper cites Soft best-of-n sampling for model alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Soft best-of-n sampling for model alignment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:06.250074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:53.904911Z digest=sha256:a2583f306fa48950007fe5aae9f45d5867ea187d3ebe60ed71774fa0444d7482

Observation 3500e86d-7259-4f5b-85ee-db9d1d77baf2 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Reward model ensembles help mitigate overoptimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:06.003667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:54.001136Z digest=sha256:1f649f4118560e1fce03fb11bb233f2d70ca4d5a8e098e5e79d5433995d3492b

Observation 8d3a0bcf-b5a8-421e-8c57-287a77f28537 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.102931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.102931Z digest=sha256:abe8fec0a09bae033f6d0a36a8c11d5370c7dbfef1642988f24c91d7e8ce5e65

Observation 5f2a30f0-326f-48cf-a273-7622b5ece621 · outbound

This paper cites Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:05.752450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:54.224393Z digest=sha256:dc8d3666923c737fc652f4a37732de468ea42699d140cd193148e14e78d3cbd7

Observation 3366544c-778c-497b-a6c4-f82a7b6c2df9 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.348620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.348620Z digest=sha256:69b61cd1b99c527ccc2dfd2611ad9178844c8295b838473548237776d59d0820

Observation 186afac0-7b57-4d07-a848-7ac07c1fadee · outbound

This paper cites A framework for few-shot language model evaluation.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A framework for few-shot language model evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.469265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.469265Z digest=sha256:5b3284c5b6e1a344427c9850cb497f645ff5cd59108619e3bdcf75947ae09361

Observation 6fdeb58c-e912-4005-876e-3cb8d2b7e924 · outbound

This paper cites Scaling laws for reward model overoptimization.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Scaling laws for reward model overoptimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:05.543335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:54.570470Z digest=sha256:0b7987901eda83138c9eb72b823297f93836a4f8792bf13ec890038c21b5249b

Observation d9d32290-4ce1-49f2-a2fe-9067a5ef89d8 · outbound

This paper cites Guided Speculative Inference for Efficient Test-Time Alignment of LLMs.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.698080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.698080Z digest=sha256:66ac551016fba5f25397c64cf645325bbd495b54f7004658df864e909e5223af

Observation ab799296-c1a3-44d0-b8c4-81e837b32367 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.801422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.801422Z digest=sha256:59462bafb03ba477bf5c37421cde303933566bfaa10636b1ad93449032e5bdda

Observation 539fb220-9f5e-44d2-8be8-54fe107213b2 · outbound

This paper cites Statistical theory of extreme values and some practical applications.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Statistical theory of extreme values and some practical applications

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:05.217536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:54.941263Z digest=sha256:04b0e3aa9836e464e4e388588acc91119e20480637af0f5259e861230bac75f8

Observation 64402ae6-45dc-4991-90ef-4421c7277eb4 · outbound

This paper cites Hilton, P.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Hilton, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:04.947535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:55.069321Z digest=sha256:2b6ada6d2dc04338be39d346161c39e753909dbfe38d40a1d29c1ec1f80b6d7b

Observation b1f6fedc-0060-4073-84ae-6bea47ec2d06 · outbound

This paper cites Risk-sensitive markov decision processes.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Risk-sensitive markov decision processes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:04.693728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:55.145296Z digest=sha256:50944bf0a63531ab7166b57382571a619ede6a0d4732993f28819e2af284b767

Observation c363c589-ea24-4b53-b71c-b72722679d1d · outbound

This paper cites Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.252415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.252415Z digest=sha256:7728470123577ef7e01d4be1151317cdc58b763fa286e53566fd25683b53287e

Observation 71faf483-00ac-4258-909e-e4c8a15e3999 · outbound

This paper cites Best-of-N Jailbreaking.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Best-of-N Jailbreaking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.320762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.320762Z digest=sha256:bf8930a9218b7f9bff2019c3a0e03c4ddf3781c5524717c91d0e1b274aca2b2c

Observation 9389ae74-a7e5-4806-a549-42a4c80e1ee8 · outbound

This paper cites Evaluation of best-of-n sampling strategies for language model alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Evaluation of best-of-n sampling strategies for language model alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:04.478293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:55.411403Z digest=sha256:b5e9c92b9809359d9624e20faf9c2c68e4ae86302b347c398059dcb37ab1633c

Observation 5e1823eb-b27b-49b4-aef6-df26a2505cc7 · outbound

This paper cites Regularized best-of-n sampling to mitigate reward hacking for language model alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Regularized best-of-n sampling to mitigate reward hacking for language model alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:04.252861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:55.545672Z digest=sha256:c94beeb7d0acacab009bdf6be2a263beed6a755753aa0fd97848584e74862ede

Observation 6ed88f36-d430-445e-b00c-1c10a059d193 · outbound

This paper cites Inference-time reward hacking in large language models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Inference-time reward hacking in large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.657057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.657057Z digest=sha256:dc2972ac28a9b7e4f7a84e6b50655c47c9b56b99b312aeb1d3d32c8f5a59136e

Observation a3223f1b-673c-43ab-ae62-693610b0990b · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.771105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.771105Z digest=sha256:a6649c0f66ffc640f3567b0605dec4a099881655fe773254b0ba9265c9cdb73d

Observation 05b1f509-175e-4ba2-8111-9c46923220e3 · outbound

This paper cites On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:03.984079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:55.916902Z digest=sha256:78ad18d5b0b187fecdd03ab8c42109bfff9e4d0648859b97e1c0a0cc3cb5f977

Observation 93333f7d-8616-41bb-b553-deae042a3f7d · outbound

This paper cites RL with KL penalties is better viewed as Bayesian inference.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis RL with KL penalties is better viewed as Bayesian inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:56.065982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:56.065982Z digest=sha256:0b7e781fd94f4cd55a92417a0f0a9f9c10fc392775a6b0bcd784430ac3041be5

Observation 4880f3a2-9ccf-4207-a706-833ace878a9d · outbound

This paper cites A new penalty function method for constrained minimization.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis A new penalty function method for constrained minimization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:03.773935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:56.187855Z digest=sha256:802f9b3a00f9c3118b94328cfe9ba067f254277f98bae06278d380c2b20b3b9d

Observation 310532ee-2c97-4b8b-af17-a07f09c7015f · outbound

This paper cites Unveiling Safety Vulnerabilities of Large Language Models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Unveiling Safety Vulnerabilities of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:56.310852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:56.310852Z digest=sha256:cb44d9dbce3ed6a1c8e186398ac474a315cc1fa7e205d74e94023387903f8691

Observation 53e330cf-3069-4bea-8c0c-a36497ded7ec · outbound

This paper cites On tilted losses in machine learning: Theory and applications.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On tilted losses in machine learning: Theory and applications

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:03.551967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:56.384645Z digest=sha256:42694cae2cbeb6d95341711a5f2c1793b9c9ebda60f340045384c49c479bd456

Observation 6dd95ac2-a8a0-4e63-bd5d-56f1633c3e6c · outbound

This paper cites Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:29:00.463455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:56.511550Z digest=sha256:730fe604b022029a48b4ae2efb6e3e5e0e9d039103dc423875435e9a374ad9a7

Observation 0642541b-daa0-41ce-8282-c91540f41148 · outbound

This paper cites Information Theoretic Guarantees For Policy Alignment In Large Language Models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Information Theoretic Guarantees For Policy Alignment In Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:29:00.250224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:56.648495Z digest=sha256:1a05b4c9d02d1b7891b0eaa84f04f5f9a925af5941223f9dbf04b00cb15b4b9a

Observation 2be4de91-739a-4700-bbfc-75b639fe3726 · outbound

This paper cites Controlled decoding from language models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Controlled decoding from language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:03.298681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:56.751460Z digest=sha256:f857308d00eecd8fb6ff497bdc991d1f02b5aa1b3808a44ec7d7c3f3448f9b64

Observation 76b71c17-22b7-47fa-a71a-192e9cdd5cb0 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis WebGPT: Browser-assisted question-answering with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:56.899965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:56.899965Z digest=sha256:48186ac6d99a4173ed3ab410aeccddfc77ad36a58e03c34f49e115de4d93e3e2

Observation b786f6be-7028-450b-a487-f725761866db · outbound

This paper cites 2 OLMo 2 Furious.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis 2 OLMo 2 Furious

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.016499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.016499Z digest=sha256:2544659e1d69dfebac7b3d682e57e3c4e3e2cafd35eb4ea7e162029d1e9acd77

Observation 727ea2c9-e0d9-49a4-a3b4-07f05e4e6421 · outbound

This paper cites Training language models to follow instructions with human feedback.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.146706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.146706Z digest=sha256:7ee05d7fab9b67891721284dfcba5cbe9ea462cdefe1c5005da70cc973c6c9f0

Observation b5753a7e-95fc-4b31-ba63-0362f90f5ff4 · outbound

This paper cites On solving large-scale finite minimax problems using exponential smoothing.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis On solving large-scale finite minimax problems using exponential smoothing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:03.055142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:57.243847Z digest=sha256:c168475cf6ad2dd5e48463490f023589e31a28001af409c858dca18bbc93a7a1

Observation 4d635e82-6652-41e2-8114-782abf0743bf · outbound

This paper cites Information theory: From coding to learning, 2022.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Information theory: From coding to learning, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:02.826144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:57.401914Z digest=sha256:9ff347dd8361b39204c4ab8a60d191a321b5abfc31b546feb08cbe99151d60ac

Observation 380b06d1-6599-4930-a23b-5248a0d52bd6 · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.548449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.548449Z digest=sha256:a595be1d5913fa6f5cdabc75d19bba9143c2ae3e8c6b723412f53964c9341011

Observation a1f2b118-520e-4cbd-9198-e42a0ac0638e · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.640777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.640777Z digest=sha256:3eab4b1ad356f2f742cada9d186ac6917e2a51775953d452ad30123f4143ef7a

Observation 1082371d-7dc0-4884-9d76-215b731ec668 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis BOND: Aligning LLMs with Best-of-N Distillation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.750523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.750523Z digest=sha256:02c5235f9abcdb7932d422fa8f2192997e4d59389a2b794bd2871a960e5f0437

Observation da09c6cb-4cdc-4fc2-b991-3d8e071a0571 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:57.885691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:57.885691Z digest=sha256:378915680c10b8577da1db20805988f234d9dc090d77ff0f4a2828791ad8ffb2

Observation 3eff8743-70ea-4993-a1df-8e259b7fba33 · outbound

This paper cites The importance of online data: Understanding preference fine-tuning via coverage.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis The importance of online data: Understanding preference fine-tuning via coverage

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:02.564388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:58.017561Z digest=sha256:2e5869b8993d4a443a7be32ff51c2fbb5636fb65fcc7ea0f84aa9c849c5aa722

Observation 43b2d2e0-e4b5-44d8-8694-8a8679a424a0 · outbound

This paper cites Learning to summarize with human feedback.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Learning to summarize with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:58.112400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:58.112400Z digest=sha256:2bbd8c0b053d19f4280602a5a2a2e21ec4792856c2e5d58b1f0bba793f0b978c

Observation b0da765a-9292-43ce-8a78-0d46ac347b0e · outbound

This paper cites Inference scaling f-laws: The limits of llm resampling with imperfect verifiers.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Inference scaling f-laws: The limits of llm resampling with imperfect verifiers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:58.222281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:58.222281Z digest=sha256:a6dcea1763bcc31f0c1f0f087bb2ec124bf46e9074dd0592dcece1f14e1d9e57

Observation 9c54f448-de15-4244-b20a-a49a0b647a27 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Fast Best-of-N Decoding via Speculative Rejection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:58.356008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:58.356008Z digest=sha256:5cbd969a6ed47504c7eaac689bf2e2957a59cd5bc63a8392d418ca9ddbbd3da2

Observation 4a49e034-e132-4766-87c8-61fa95b9cdb6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:58.458027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:58.458027Z digest=sha256:7975ee4006a4f522294add75dd56e9d8a21ae5b5f3e10b80ce3b0d4fb788d51e

Observation 0e3eca5e-52af-4aa6-970f-3d4ee506d061 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:58.572765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:58.572765Z digest=sha256:7756abce5ea759d654231c93d2a564e6a08286c1d9768ee804d346fad3f21acd

Observation 55bbc992-aee4-4daa-b6ee-a5cf6912e674 · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:02.304528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:58.704734Z digest=sha256:f35ecbb4e6b6ea11d7dd77f0363d13be40bc31c06dbdce5b656f15058eaffb1f

Observation db710494-7c25-499f-9fda-597c279565c2 · outbound

This paper cites Robust variable selection with exponential squared loss.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Robust variable selection with exponential squared loss

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:01.995933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:58.851458Z digest=sha256:115d32965844b5deb31cdb149f346ea0510871e2a06b1b87d79383050438f42a

Observation 12b2e296-184b-436e-b81b-7bece514809b · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:01.699260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:58.937433Z digest=sha256:0073a641c55a7564e09a74727b89b3c2935d7c9f5af4acd2575191ddc59c1de7

Observation c9134f7e-168b-4b3b-8458-985d2c47c02d · outbound

This paper cites Asymptotics of language model alignment.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Asymptotics of language model alignment

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:01.364536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:59.076528Z digest=sha256:ca0e9a5657cfdf6f0cc9008d59f26b64cf5ebaeb30642b0ddca941073669c5e9

Observation 093679e0-110d-4480-907f-5811acb13c66 · outbound

This paper cites Convergence of the inexact langevin algorithm and score-based generative models in kl divergence.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Convergence of the inexact langevin algorithm and score-based generative models in kl divergence

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.200479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.200479Z digest=sha256:3e0a207cbd77f60057d8492551ae381c8da4e8405f0da3886da94599f98fc0e3

Observation 683f5707-3ded-4ea2-9359-4f792525dc08 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.337652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.337652Z digest=sha256:279a4e6723b7c77b84aeecf9cd7e8cc3b17df34159e0614d02d199b031eab92a

Observation faa4c54d-4083-4620-95f4-cb0d669efcad · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.418783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.418783Z digest=sha256:ef49ca8c53caee8036e987d97aa18c4637cba27682f7c3219a7f923986686ab4

Observation c246fe7b-c98b-49c4-ab0f-333531e32621 · outbound

This paper cites Sharp Analysis for KL-Regularized Contextual Bandits and RLHF.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.510021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.510021Z digest=sha256:670d192d59488072d1b2e9f60d3d609144afa642c24aeb0041e95f8c642c47f4

Observation 6d617809-b2c2-46f8-acdb-1da6b3ccd165 · outbound

This paper cites Calibrating sequence likelihood improves conditional language generation.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Calibrating sequence likelihood improves conditional language generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:29:01.134347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:28:59.653529Z digest=sha256:94a2833cad2c6070e0540adcc60c4803670d91b468e40f80f7ef6d8399eeb223

Observation c3eea33b-c332-4ce8-a9da-7bb7bec78699 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.743545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.743545Z digest=sha256:7f026e94faa45f88bd683b5b15c1edf24a59acae71169dc00b1a5f54fb623235

Pith citing papers

Observation 3545a9e9-d152-4a23-b9c4-7580a44a94cb · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.637485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.637485Z digest=sha256:314ef2cb10b2f80379c68a2d29ad774d6849915b58d328eb0b556be7f4ec2612

Observation 8584fc7a-97be-49fe-a67a-187801539ab3 · inbound

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment cites this paper.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:29.247752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:419505d9fae0ed0b3dfe780048325832587ccf11238e0c9eb920aa837ec277a7