Pith. sign in

Paper Citation Record · LEDGER

Soft Best-of-n Sampling for Model Alignment

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2505.03156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03156 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:05:48.302057Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:04.163654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:21:09.732556Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 119fca26-74da-4941-8c97-a54743aebe9a · outbound

This paper cites Cooper- ative inverse reinforcement learning,.

Soft Best-of-n Sampling for Model Alignment Cooper- ative inverse reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.081029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.081029Z digest=sha256:33c824e532a84ba9a91c99966598fa369424a6b57b7edc7ca68d6f355a626de7

Observation cf92ce88-b690-420d-b35b-72c96389e60b · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Soft Best-of-n Sampling for Model Alignment Scalable agent alignment via reward modeling: a research direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.085299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.085299Z digest=sha256:40c56de70bf4828006c7c3b430120394c6a193045b38afad2fcb5921fb024558

Observation 249298fd-ba9e-4eb5-8b51-f976b462a6fc · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big?.

Soft Best-of-n Sampling for Model Alignment On the dangers of stochastic parrots: Can language models be too big?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.089813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.089813Z digest=sha256:d9b4431e5dae412a60104f73bdb13bfd9024d0dfa8895be8dc2a7bb2fdceca82

Observation 9e7e46d9-848c-4a7e-b168-af1e14fe537d · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Soft Best-of-n Sampling for Model Alignment On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.093738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.093738Z digest=sha256:012cab28fe61f67e6de85bf666ede3477f2ef1d6a0b94c57fa07cbbd3a6393ef

Observation 946bc06f-d098-4f40-87dc-0468fef3645b · outbound

This paper cites Information projections revisited,.

Soft Best-of-n Sampling for Model Alignment Information projections revisited,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:49.050354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.098398Z digest=sha256:15bdad3755d296f2cb1db4e66733bf6d455a1d7557644572b78a2f0c8d8a3432

Observation f283eacc-5d4a-479e-8d21-2eddd2dda446 · outbound

This paper cites I-divergence geometry of probability distributions and minimization problems,.

Soft Best-of-n Sampling for Model Alignment I-divergence geometry of probability distributions and minimization problems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:49.038436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.102343Z digest=sha256:79bcd33e89d3e90bfe066f843099cdf1925ad5a78504657aadad4cde1354b16e

Observation 7f67f304-daf5-4c85-bda1-dd5cd54db855 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Soft Best-of-n Sampling for Model Alignment Training language models to follow instructions with human feedback,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.107841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.107841Z digest=sha256:c363428165d22c578501bede594b98071767acfcd82afa5b962e46e1700a7963

Observation 3a2e02ca-a36f-49a0-a737-b440876d2dc9 · outbound

This paper cites Nonsymmetrical distance between probability distribu- tions, entropy and the theorem of pythagoras,.

Soft Best-of-n Sampling for Model Alignment Nonsymmetrical distance between probability distribu- tions, entropy and the theorem of pythagoras,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:49.013830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.111837Z digest=sha256:4a8bc944ac6b7c93ff8e368574498b865060882090ae3c860ae94be343d180a2

Observation 8a28a2b8-2688-4687-8c5e-b7e62b195914 · outbound

This paper cites Generalized projections for non-negative functions,.

Soft Best-of-n Sampling for Model Alignment Generalized projections for non-negative functions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:49.002498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.115719Z digest=sha256:96532b2286bd2f5fe6e9d07717c0fbf5b4548df58445c7413c7ab84d5851a590

Observation f5c29976-c02d-4c29-8caa-b97eeecacc42 · outbound

This paper cites Model projection: Theory and applications to fair machine learning,.

Soft Best-of-n Sampling for Model Alignment Model projection: Theory and applications to fair machine learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.992085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.119692Z digest=sha256:9f638b85a329dd19723b1fdfe1acfaef11da35879e59b4d12a4b2fe30633e169

Observation b771202b-5e63-4a7c-bbb1-f079d7d70ccb · outbound

This paper cites Efficient methods for generating some exponentially tilted random variates,.

Soft Best-of-n Sampling for Model Alignment Efficient methods for generating some exponentially tilted random variates,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.981486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.123891Z digest=sha256:c2fff572362d1a9fbcd47876905e587b295d177e19b22208e8a35ba35fe316b5

Observation 75cf8fb3-1b08-4dd0-945f-178560738c8b · outbound

This paper cites Sampling exponentially tilted stable distributions,.

Soft Best-of-n Sampling for Model Alignment Sampling exponentially tilted stable distributions,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.969362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.128252Z digest=sha256:9abbed0e690971f6dbbe069d931bbe4f3ac5e7c123cf3e70bfb28822b5efd976

Observation 18c1adb0-c238-4dba-ae2c-cd710a37a1bb · outbound

This paper cites Efficient exponential tilting with applications,.

Soft Best-of-n Sampling for Model Alignment Efficient exponential tilting with applications,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.957333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.132190Z digest=sha256:7aa184de21d3a362a5ce7150a9db761e63aa80f720b758fb894106d375d6dba0

Observation 3f0cc25c-8953-46a5-a74b-ac4a2b5b1723 · outbound

This paper cites Information theoretic approaches to inference in moment condition models,.

Soft Best-of-n Sampling for Model Alignment Information theoretic approaches to inference in moment condition models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.945263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.136501Z digest=sha256:70f435fc66188be9c47698eb5beacccb18b3145de10f9cc27223f3331bf3c965

Observation 2e1e1ce8-fe33-4302-b866-9e1cf4f21455 · outbound

This paper cites An information-theoretic alternative to generalized method of moments estimation,.

Soft Best-of-n Sampling for Model Alignment An information-theoretic alternative to generalized method of moments estimation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.931763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.140626Z digest=sha256:92078e7d41c9ae1f0beeaa646e7d151b385572c627633fa5fe59ec3a03ffc16d

Observation f1b6ac90-7687-402b-9e7b-0ecd0cb01d3d · outbound

This paper cites On large deviations theory and asymptotically efficient monte carlo estimation,.

Soft Best-of-n Sampling for Model Alignment On large deviations theory and asymptotically efficient monte carlo estimation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.917490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.145308Z digest=sha256:05f796a2e40ed2b5b5e62838b54d2f4e733a570ba97715627b2020d91dd591f2

Observation c621a5d3-a6f6-4a09-ac57-3fc44a9ab953 · outbound

This paper cites Importance sampling in the monte carlo study of sequential tests,.

Soft Best-of-n Sampling for Model Alignment Importance sampling in the monte carlo study of sequential tests,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.904297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.149880Z digest=sha256:fdb473a37b6f6a9c1187448a1cf80266b699c7cd1f837abf90ee415b7cb30447

Observation a74c6439-fa4b-461d-aee1-3c9a06712ff1 · outbound

This paper cites Asmussen and P.

Soft Best-of-n Sampling for Model Alignment Asmussen and P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.788331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.161423Z digest=sha256:15f9a5022fb56138c2ba410ae95205b4fb57fcc3f92457aa1cd784a620847164

Observation e89ab501-f4f2-460e-9aa0-b3a765d73e5f · outbound

This paper cites Deep reinforcement learning from human preferences,.

Soft Best-of-n Sampling for Model Alignment Deep reinforcement learning from human preferences,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.166674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.166674Z digest=sha256:585aee4314107c00d7826770a78b4a95e2cf411404d15505d9460f113050debe

Observation 6423f38b-30b5-428b-98d9-516555a932ae · outbound

This paper cites Learning to summarize with human feedback,.

Soft Best-of-n Sampling for Model Alignment Learning to summarize with human feedback,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.769646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.170857Z digest=sha256:e979ed979e808b6613628395d7c4f22f6364c4b29f31b46c0e9d37adeaf9d37e

Observation 9f750393-b763-420c-b82f-4d8fbc8f859c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Soft Best-of-n Sampling for Model Alignment Direct preference optimization: Your language model is secretly a reward model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.174729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.174729Z digest=sha256:0716cc6aa4e3f015bc6bf823b09a761268541a8ec55be0e52c6404001a62bb79

Observation 8db3ae53-eaec-4bb5-85a3-9165a4ab76e3 · outbound

This paper cites Scaling laws for reward model overoptimization,.

Soft Best-of-n Sampling for Model Alignment Scaling laws for reward model overoptimization,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.179101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.179101Z digest=sha256:bea83df24881c089cb3465cff430b417e6f054cf87115840ef44805a9e0c35f1

Observation 93ef3e00-e383-4770-a2c9-4eb1a08c3325 · outbound

This paper cites Controlled decoding from language models,.

Soft Best-of-n Sampling for Model Alignment Controlled decoding from language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.744388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.183426Z digest=sha256:682eb268c525a8f99d88f0652469896181a93c99f255fe6a535b108fceacdbe6

Observation d98e3fc8-7e38-45b1-b9b4-2aff231eb00f · outbound

This paper cites Asymptotics of Language Model Alignment.

Soft Best-of-n Sampling for Model Alignment Asymptotics of Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.187495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.187495Z digest=sha256:dea1a66f40021f51047d0225429a00f4f335e3c222d543924051e4be055e0246

Observation 04f7c2cc-c104-4a9d-96cb-063413f4a91e · outbound

This paper cites Information Theoretic Guarantees For Policy Alignment In Large Language Models.

Soft Best-of-n Sampling for Model Alignment Information Theoretic Guarantees For Policy Alignment In Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.191821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.191821Z digest=sha256:58809562973585f72d294b45d329e20a25ae966aed08f1114108aaafc0258d00

Observation f24ea7bd-375c-4c7e-b546-6f6951671e76 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Soft Best-of-n Sampling for Model Alignment Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.197392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.197392Z digest=sha256:c97044fba419d33d8c9fb80bbafca4d58a125a072b0f9d2c3f618b8c6d0613ba

Observation c2fc834c-8175-432b-a92b-da2c6f4c1fbe · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling,.

Soft Best-of-n Sampling for Model Alignment BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.718686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.202277Z digest=sha256:e8166be33a045c364e8bcd125370514f95309bdee852c8a4cde4602ec4879840

Observation ab98e271-b6aa-456b-81fc-beb9909ec602 · outbound

This paper cites Regularized best-of-n sampling to mitigate reward hacking for language model alignment,.

Soft Best-of-n Sampling for Model Alignment Regularized best-of-n sampling to mitigate reward hacking for language model alignment,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.704839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.207169Z digest=sha256:70d83c682b46b5fd155309dbe410f4d9af28ba014ba9fada7e2527f926282584

Observation f9534c11-8bf5-4953-9c4a-504feb99b1bd · outbound

This paper cites Evaluation of best-of-n sampling strategies for language model alignment,.

Soft Best-of-n Sampling for Model Alignment Evaluation of best-of-n sampling strategies for language model alignment,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.691949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.211696Z digest=sha256:95c8dfef7ff6229d2ddadd589c8bd079e57c9a2100c65b205bdbe3539d8de0dd

Observation e02ae75a-590f-49a4-8040-c3fb7431a5a0 · outbound

This paper cites Variational Best-of-N Alignment.

Soft Best-of-n Sampling for Model Alignment Variational Best-of-N Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.216166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.216166Z digest=sha256:ab0d8bf047a896ddb062c737124feca26819dff6ae2b1989bf18f589c9d701cc

Observation 2c07a9fb-f792-4ab8-9578-9c07c7e5fe53 · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

Soft Best-of-n Sampling for Model Alignment TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.220969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.220969Z digest=sha256:f915f5d5278e21099ee7437e7b1743d410cb40d780643a7f075df1487f61a5bb

Observation 4fd2cfae-42bf-48ad-b74b-23ffc57e033d · outbound

This paper cites Accelerating best-of-n via speculative rejection,.

Soft Best-of-n Sampling for Model Alignment Accelerating best-of-n via speculative rejection,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.681151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.225707Z digest=sha256:357dff75287ebecdacdc87aab25ed79664908e465227e36f43534627f707dc71

Observation d0fdd318-2ce1-4c20-a124-eb5ecea956c1 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Soft Best-of-n Sampling for Model Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.230989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.230989Z digest=sha256:5f9855a65844bfbe84ce2694f81f04f5765c91d2ecd2f5c6571e105f934d86ff

Observation c645bc33-49a6-47c7-b575-cfb3503b55cc · outbound

This paper cites Measuring goodhart’s law: Towards an evaluation framework for open-ended generative models,.

Soft Best-of-n Sampling for Model Alignment Measuring goodhart’s law: Towards an evaluation framework for open-ended generative models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.669786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.236594Z digest=sha256:1d63a069281dae29c70a8767ffc4e5117efc124bad7e6edd06d24826c69f0ea5

Observation ceafe069-1a87-476d-b614-6390ee840220 · outbound

This paper cites Theoretical guarantees on the best-of-n alignment policy.

Soft Best-of-n Sampling for Model Alignment Theoretical guarantees on the best-of-n alignment policy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.242261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.242261Z digest=sha256:cde0e92826cbe1f7e3632a15a1703c5084c84bd8f02be25b67855fafc91f11ca

Observation 892bf843-84f8-4fe0-9515-8425ca4ae081 · outbound

This paper cites A better bound on the variance,.

Soft Best-of-n Sampling for Model Alignment A better bound on the variance,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.658260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.248193Z digest=sha256:0a5a95f790ff936987b1dec2a8a1e6f846dfc7f8419ca809607e35a4990e7907

Observation 533c34ef-ddc7-4d3f-ac02-7420a99a5f4a · outbound

This paper cites The accuracy of the gaussian approximation to the sum of independent variates,.

Soft Best-of-n Sampling for Model Alignment The accuracy of the gaussian approximation to the sum of independent variates,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.647257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.254243Z digest=sha256:878d18cbef1fe2b5cccb6f14bca5f000e3232651ae9f51f807f2f9dbafb38140

Observation 81e65984-29a4-478e-8ba7-0e9960761ac6 · outbound

This paper cites Esseen, On the Liapounoff Limit of Error in the Theory of Probability.

Soft Best-of-n Sampling for Model Alignment Esseen, On the Liapounoff Limit of Error in the Theory of Probability

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.259251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.259251Z digest=sha256:6d6a37aab102bd590cd9e8ae67fa3dd690858519be64bf027ed8ca3a6229b79d

Observation 7d1abfc7-a040-4f29-8dcf-f458f7d64073 · outbound

This paper cites an unresolved cited work.

Soft Best-of-n Sampling for Model Alignment Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.264764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.264764Z digest=sha256:f21b0fb8bc70e07a37b9b62a96bb61620ddf2ce86c604664143977523a81e83d

Observation 083c57c9-863a-4c12-9ff8-8ebcbe127407 · outbound

This paper cites Concentration inequalities and martingale in- equalities: a survey,.

Soft Best-of-n Sampling for Model Alignment Concentration inequalities and martingale in- equalities: a survey,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.621781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.268838Z digest=sha256:af7cebdbb5a8035b7f1c6b75bde874dd6b6753149a16bba809af8cdf83d685a6

Observation b9c04015-e82c-4d71-9d7a-194f7daeaeb7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Soft Best-of-n Sampling for Model Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.272938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.272938Z digest=sha256:71b1af15dd3e47e60259240608b14e2981395c90234ae9ba5f4be79697809ef1

Observation 78b093f8-b8f5-447d-8501-fd2176405866 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Soft Best-of-n Sampling for Model Alignment BOND: Aligning LLMs with Best-of-N Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.277881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.277881Z digest=sha256:f5a76b13a7932f030f61d5458179c68441a2fbdc6e2e761e32531d0961a628f5

Observation a50f8f4d-62e4-4be6-b368-8a201baf145a · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Soft Best-of-n Sampling for Model Alignment Chain-of-thought prompting elicits reasoning in large language models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.283589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.283589Z digest=sha256:12f589056c99612458fe34adccbf64b0bf85ef1e96ef8ddd463d1de89b2719ce

Observation cb049b0d-40dc-487f-b0e2-e9a2ab5c84e9 · outbound

This paper cites Let’s verify step by step,.

Soft Best-of-n Sampling for Model Alignment Let’s verify step by step,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.598969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.288183Z digest=sha256:97f5ef0b40bf23ca608ee6cd088fc3d3af7c56b5d8d1b5a028879f2ead5f915a

Observation 5c7c80ab-4c03-4571-b512-513e03914768 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Soft Best-of-n Sampling for Model Alignment Solving math word problems with process- and outcome-based feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.292807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.292807Z digest=sha256:6b7e35242e270901244559de61e3022885aac81a1238e4bfe3c8a6ae40ed3c2b

Observation c75a4f7a-0f72-48d6-84ff-dbe03275e144 · outbound

This paper cites Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods.

Soft Best-of-n Sampling for Model Alignment Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.297295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.297295Z digest=sha256:51e5ec26900d01985709a68c9197bd9f200965f646ada9ae7c84db9e2d3842aa

Observation 8f8b0539-0aa3-4336-951e-5471eb9bfe19 · outbound

This paper cites symbolwise.

Soft Best-of-n Sampling for Model Alignment symbolwise

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:05:48.587325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:05:48.302057Z digest=sha256:ef3fd7b2e25e67419c8930d7e24405d997bdafdc7dd42483182bdb5f4ac50583

Observation ef5650b5-5685-46a2-adef-a69faba3f5e2 · outbound

This paper cites Available: https://projecteuclid.org/euclid.aos/1176343541.

Soft Best-of-n Sampling for Model Alignment Available: https://projecteuclid.org/euclid.aos/1176343541

Reference 1976

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:48.154989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:48.154989Z digest=sha256:2dc41a6f114d8c3025d6fd56df416c73135cb7db737235cd923a6310085d2c2d

Pith citing papers

Observation b4018f26-4601-459e-a03d-bbf657a72bf4 · inbound

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology cites this paper.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Soft Best-of-n Sampling for Model Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.298599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.298599Z digest=sha256:4c596b11af05a562ba9114911bf104eb439926fc7178a8f1f517ae603ad59989

Observation 709c5b54-ebe8-4c6f-b65e-954fa42b7bde · inbound

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment cites this paper.

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment Soft Best-of-n Sampling for Model Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T23:54:26.281624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:54:26.281624Z digest=sha256:90377178ce872bf94d219d5202529a9470ea416e2327a2438a26f22b8c52e257

Observation a7e70538-ab06-4cfb-a7f6-7b697f5e8681 · inbound

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data cites this paper.

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data Soft Best-of-n Sampling for Model Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:09.735874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T09:27:46.993773Z digest=sha256:cff20a5105e5c0dee08f9a018eaef740379c0a42456d7913599bba2dd64207fc

Observation 8c0b00f4-c033-4652-ab2c-fe81225ccd6e · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Soft Best-of-n Sampling for Model Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.318110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:0afb50078bcaf5fac6ba6ca4af64c45f4c67958179dffa2394fda1f077fa385c

Observation ef5d32b2-c04c-47a7-b003-7bfa0784736e · inbound

Theoretical Limits of Language Model Alignment cites this paper.

Theoretical Limits of Language Model Alignment Soft Best-of-n Sampling for Model Alignment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:56.867358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:18:37.614335Z digest=sha256:1d0f28a6dece3cbcd56499931213bf1473d92629eb5eecbdccddce69dc6b5757

Observation 9ed9df29-4211-437d-97f5-2d7e42b33539 · inbound

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning cites this paper.

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning Soft Best-of-n Sampling for Model Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T02:23:04.515705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:23:04.515705Z digest=sha256:e3310fd14e9f2d423f57cf9a1779fa209fda013a6ef0887f8fc41ff39bff8ce6

Observation c56500bb-cc81-4324-9519-2830e00b443c · inbound

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility cites this paper.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Soft Best-of-n Sampling for Model Alignment

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.163654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.163654Z digest=sha256:2151bd2226da9131e9e437c4e987e46158c1ddf3f97fdd4b77313f732708c6c9