Pith. sign in

Paper Citation Record · LEDGER

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds

As of 10 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23673 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:45.634287Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbb2eea2-eefa-4813-a6ce-abe16523ff2b · outbound

This paper cites write newline.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:39.878545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:39.878545Z digest=sha256:b1999046004bf819740a06d868c627179b45d9eaba20cdc0a2599b1217c86695

Observation f25d9bbb-539f-4ffa-a943-5380cdd1e2ee · outbound

This paper cites Online learning for linearly parametrized control problems.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Online learning for linearly parametrized control problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:57.055986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:39.976257Z digest=sha256:6575606cde58bbe0b9981b1410d7b5e296bf59ee61bdbec5ec86c2727c19c461

Observation f63f9162-3596-4bc2-b4e7-b14c5fb6ed7d · outbound

This paper cites Reducing dueling bandits to cardinal bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Reducing dueling bandits to cardinal bandits

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.865663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.043920Z digest=sha256:34ad241c81103bb5d4648d873f206a92a7ea7a07d4df349a2d8c4832a87f5ccf

Observation c916829b-9808-4ebb-a0ce-b4bda5cf7a20 · outbound

This paper cites S., Hu, W., Li, Z., Salakhutdinov, R.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds S., Hu, W., Li, Z., Salakhutdinov, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.138320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.138320Z digest=sha256:cff83c6700c6f3140b3d90228eeaa6a5a8e6a6c9a34994a19b1e2a2dd9b9256c

Observation 661e7698-f079-457d-91a0-dcbd759a411c · outbound

This paper cites R., Daulton, S., Letham, B., Wilson, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Daulton, S., Letham, B., Wilson, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.750917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.241854Z digest=sha256:e34653b43ee697aa69cd3bee893899ec2d2210f8425138bdc095026a45197fbc

Observation 3c8de591-d8e2-458b-9040-58d0cb36bd85 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based online learning with dueling bandits: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.449269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.361927Z digest=sha256:240a1276cae7332b81a787f1316dc94808ec9c0a51870445032a058103c15ba8

Observation 3b08b2bb-24a7-42e8-af6a-df979b24cc66 · outbound

This paper cites Stochastic contextual dueling bandits under linear stochastic transitivity models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Stochastic contextual dueling bandits under linear stochastic transitivity models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.187735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.442785Z digest=sha256:287709404bcb1276a1cd2db7c875457f1579df3ef1b773f7991a2ec47d560dd3

Observation c5f9d829-ad58-45fd-b147-a251fe5e750a · outbound

This paper cites Mat \'e rn gaussian processes on riemannian manifolds.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mat \'e rn gaussian processes on riemannian manifolds

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.922717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.521123Z digest=sha256:5cde2d391dac51c6fa99af1f4e4c5b124a2cfcc34ffea8e773ec69642a9e7f41

Observation 1cbd7aa1-5dfe-4a1d-9f66-859d3c207769 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.628781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.628781Z digest=sha256:68570988f1fc2c61399f994ade7bdba04c47e59905205e07828bf3a8d586745d

Observation d1b04aac-25ff-468c-a4f1-279ce8a31e28 · outbound

This paper cites A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.734979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.734979Z digest=sha256:578ec8228f338584eee3e7e553b749b6775c7d9b1b3f1e58299a95c248ff3663

Observation eb9496c5-f672-4075-a016-e37658ba9289 · outbound

This paper cites Instructzero: Efficient instruction optimization for black-box large language models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Instructzero: Efficient instruction optimization for black-box large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.729489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.835746Z digest=sha256:eae7e1008ae4847662abaeea16f23f270ef6936ab7bf2370785a84bf4be51879

Observation 72f068b0-e951-4dec-aabc-ef40b7e314d5 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.901099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.901099Z digest=sha256:be0227e60ac282e40937c5dcf7181d5871de551e1412298e6aad09beedb93142

Observation cd4f5627-73e6-42a9-9a40-4be5c0d094e4 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:55.354090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.994901Z digest=sha256:08110aa1098f9e6dc822cad1173dcd8bb21533a5282b98451b7d371e467167af

Observation 3a2067e6-4446-4226-8812-ca1d198dd0cc · outbound

This paper cites and Steinwart, I.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Steinwart, I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.009181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.069695Z digest=sha256:1543e419cbf4e5ef8858b2cd57e35aa5b7047c41d2cc3bee3b620c500c8ee20f

Observation bf1ae0a8-8f2e-4f09-9963-fb340e6a3851 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.743049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.168084Z digest=sha256:f08f02177ad4920f9acb242192bdd7150e5e6b1b5abb3b364bb046e7d0ca2cfc

Observation 6dd47c75-eda3-44a6-93c1-9020fbd86d56 · outbound

This paper cites E., Slivkins, A., and Zoghi, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds E., Slivkins, A., and Zoghi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:54.433457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.269844Z digest=sha256:7b361858b45f029bfaa536c145600154b707976eb5100803320449e2d4e96bda

Observation 279f94d4-b194-4864-b290-d74173a76ca7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.070394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.351047Z digest=sha256:775c3a0877cca33acf8df6bfcd861f34d185f0ef89b548501cf8a724291e40b6

Observation 90cc3a77-2405-4f42-bf4c-0d4f49d85df7 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Improved optimistic algorithms for logistic bandits

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.772960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.453532Z digest=sha256:bf82a9e4c34eb068b978d2a4626d6992ac60d7bab98ed964c5bae91e8473c9bf

Observation 3fd1dca9-b811-4db7-8ba6-edaa73054cf9 · outbound

This paper cites A Tutorial on Bayesian Optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:41.544682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:41.544682Z digest=sha256:9adfdefaaece83817967f6b2fdfcd58b8416ea5ac1e1f937ca83de88f2338229

Observation 18fce1e8-13d2-40ae-9586-1d9ec0b24e46 · outbound

This paper cites R., Pleiss, G., Bindel, D., Weinberger, K.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Pleiss, G., Bindel, D., Weinberger, K

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.480243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.638418Z digest=sha256:88bea3026e33b3f1592a9778f86412c85448c5917c4f7577aedcfebb9901d3c5

Observation b5164a3c-46a1-4518-9082-846da787c71c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:53.193134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.704728Z digest=sha256:09c5a410ba32d34ecd9befb1513fb4d88a71658be1660aafbe2fd429effda02c

Observation 1cc4fddb-d77a-422f-bbb7-62bba9674992 · outbound

This paper cites L., and Thomaz, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds L., and Thomaz, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.899224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.773886Z digest=sha256:bc4b14a4b5935e9c0ce91a12b75ac762810622535c7e9e7c01a1f2bcf6780836

Observation f82de9f4-ae72-47bb-a88c-9b04475e04bf · outbound

This paper cites and Yang, X.-S.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Yang, X.-S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.600042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.930302Z digest=sha256:6893b4834eccdd7c8d9a344eff1e730a8775a079afbb11287f38322041915c78

Observation dad41594-2365-4926-9a1e-5a488c4bb929 · outbound

This paper cites R., Schonlau, M., and Welch, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Schonlau, M., and Welch, W

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.054479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.054479Z digest=sha256:03f45b7e245dac12df289ad43bcf7f658dab2d7cf118c0e0914dce45ccf4f595

Observation 807202dd-e4de-4525-b1ca-f3f9eeb24320 · outbound

This paper cites Feel-good thompson sampling for contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Feel-good thompson sampling for contextual dueling bandits

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.335804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.145030Z digest=sha256:aac7a36b67d5fc9ec13d9a1aeb5f24a72d01ce1b6b2e91429be8efe863d4935a

Observation 70eed8b9-bb8e-4b6d-b810-cfc21206ee71 · outbound

This paper cites and Scarlett, J.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Scarlett, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.109171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.274206Z digest=sha256:1cbeb5948f0a41b3c12a0ec43de19aa93d829fbae3891f1f2cf8c8a7a672c197

Observation 6fa7c559-fe37-473b-916b-3147b64858b7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.870471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.368045Z digest=sha256:a59a80be4044703053fa83d23c916e833c1c0657718d77baefaa6844a5a043cb

Observation c52c829c-9edc-4326-9bd9-17ec706ccd74 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.603727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.497274Z digest=sha256:86d560e17d114452885a4304291ab7485f05049d20f48478d0564432313bfa03

Observation 3e05f0f4-440a-4c96-815a-b223275df31c · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.595875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.595875Z digest=sha256:44245a61641f63456f7f4ed2a6982d751423f6d409e4cb98d02d8053f53fa1e6

Observation 05359119-0546-450b-85ca-779f84c2d013 · outbound

This paper cites Kernelized offline contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Kernelized offline contextual dueling bandits

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.302536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.722443Z digest=sha256:233e717d8d07d88f449c6612bf4f3e9989acd122b89294b9b6a2ecbb49726c8b

Observation 5a862fc4-f9e7-4c39-a3ff-b6c87c6ebf2d · outbound

This paper cites Functions of positive and negative type, and their connection with the theory of integral equations.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Functions of positive and negative type, and their connection with the theory of integral equations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.069454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.844443Z digest=sha256:db37e5d7aace4042a5848c43ddc39bc0e7097e61d2b5e50109207c4759d981bc

Observation 1ce9876e-4bd5-4bb8-86be-301ac54aa385 · outbound

This paper cites Projective preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Projective preferential bayesian optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.884143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.936733Z digest=sha256:3fcce574fc979e088e064ee715695b69a3d7b4314d52c2edbdea05d3612e5f92

Observation 849b6473-3f9d-4ee8-883c-119846cf8701 · outbound

This paper cites Dueling posterior sampling for preference-based reinforcement learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling posterior sampling for preference-based reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.040733Z digest=sha256:af6caa9ddf7a22ddd13d4facc2e2bab5ad33c065d38c6a49f65e509f6c660737

Observation 8646b46f-3fbe-4102-89a5-ea926abb3064 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.162585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.162585Z digest=sha256:4e57054d2b8889763b07b5d37657b4e1de9afb8b12c578632e910853cc28cc21

Observation 4f7d64a3-69ac-49b3-a6b5-ab2d59de08c1 · outbound

This paper cites Bandits with preference feedback: A stackelberg game perspective.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Bandits with preference feedback: A stackelberg game perspective

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.681901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.268258Z digest=sha256:391e74d80d2fe388b54a242c977a67d7a9dee757051f009bf097ff08be94533c

Observation caa9fe46-d605-42c5-bc2a-bcebf3c6eac5 · outbound

This paper cites Scikit-learn: Machine learning in P ython.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Scikit-learn: Machine learning in P ython

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.415191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.415191Z digest=sha256:f4568ea770af02e0d2e97b9bddb90bf09d04501a677f1fca0e1dc9bef7da5055

Observation 166d52dd-32ae-4492-b81d-5df8ea1d1138 · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal algorithms for stochastic contextual preference bandits

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.504449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.528257Z digest=sha256:f3bb24bbaff46cd8b20a580ef28e60e94f5891a02e30bb0dc480c588df9c51b6

Observation 7d52213e-1a68-489b-9e55-ee72fe4bb353 · outbound

This paper cites and Krishnamurthy, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Krishnamurthy, A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.376160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.638184Z digest=sha256:cec6eb54856978e5b169bc1e65631cfdec465d5318c189ccde79f3e4a84e0cca

Observation 889bb251-af6e-403a-9647-950bd1ee4e92 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling rl: Reinforcement learning with trajectory preferences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.226512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.712472Z digest=sha256:91c084905a8339f08d211330e97f4c08a05170cbb3eb11d9a4dd1b677baa5ccd

Observation 1c144d77-19f3-46d8-b247-76c1035bb965 · outbound

This paper cites A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.071549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.782700Z digest=sha256:312fdfdedce49525ba7f61274de0d2d81ce7d7ed6d1c87f689b6a2c7469fbdf1

Observation c2c9467f-1cca-4af2-bb75-07d3faab1062 · outbound

This paper cites Lower bounds on regret for noisy gaussian process bandit optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Lower bounds on regret for noisy gaussian process bandit optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.875455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.853547Z digest=sha256:8e161362523d031e8cc65a08320d9470ad0e5f990a7f914898c9629aac33217b

Observation b123e8fe-ebfc-4bcf-940c-41a2afa6489c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.701942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.935161Z digest=sha256:5d6b1353b65aa82622efc62a2244696d41460cb702c838eba59db632d923e4c2

Observation 3758a7e0-2e92-465b-bbf7-cdeb81572b1c · outbound

This paper cites P., and De Freitas, N.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds P., and De Freitas, N

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.036902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.036902Z digest=sha256:5a7ba6e930084b757fb3c9c00735c20228e87fbd0d876b8eb4960c9f92445e1c

Observation 23b44e2a-b881-4f5f-981d-8fbcc8ea02c5 · outbound

This paper cites M., and Seeger, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds M., and Seeger, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.499940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.117889Z digest=sha256:5a872e5b9b8aeb54f639ae7b84ffe0a0579a48979d49b3ca39110bf8b3d9a493

Observation e8d65b47-20a1-4c62-9db7-57bebf343634 · outbound

This paper cites Towards practical preferential bayesian optimization with skew gaussian processes.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Towards practical preferential bayesian optimization with skew gaussian processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.311254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.202090Z digest=sha256:d1b0ca54dd91e41ceacda8b0284b1688ed8d5ae40d3ed80ed1d12044c4c3beba

Observation a985fd11-9b1a-4b58-a2eb-372ec3ca4b69 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.159137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.270455Z digest=sha256:d97f61b8643558a73dec3d6e6e3fb92cdb6dd145835afae9293dfa43cb47fd79

Observation 08434648-ed7d-40f3-bf5d-6b6810b15521 · outbound

This paper cites Optimal order simple regret for gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal order simple regret for gaussian process bandits

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.991095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.350553Z digest=sha256:764e431bc852a1c15f91994a6eae5f7fe12446f368785481aa5379dad1f7e976

Observation b9381b72-7c92-4d46-9a8c-8c9188bc71d4 · outbound

This paper cites On information gain and regret bounds in gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds On information gain and regret bounds in gaussian process bandits

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.847219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.435530Z digest=sha256:efbf4736389f61b8febce51fe9bfe210094ad5fe9523e5fe1e502190a603ce15

Observation 4d154dbc-42cf-4bea-8214-ebb4695eacbe · outbound

This paper cites Finite-time analysis of kernelised contextual bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Finite-time analysis of kernelised contextual bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.739454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.531719Z digest=sha256:af95b9a26fb48357041972dfb90a3ec9cd7dc82283fda60031ac2ee4ddb5cff3

Observation b91b39b6-17ce-4ffc-a485-0509836099e9 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.477948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.601543Z digest=sha256:50408db59b378f1706af308ffd5ff2d6bc03f0ab2fb3d0ded7cc4d1165ea23b0

Observation 4fd5c7cf-d33e-4877-9957-135aea590e2a · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds High-dimensional probability: An introduction with applications in data science, volume 47

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.680747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.680747Z digest=sha256:c3668f0f358bde49d3910ae936863c61e1df3bd123b2794c5fd5398a11c00dfc

Observation 25dd1c45-fc65-4544-b0e1-c48f8c1ee64d · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.210392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.789087Z digest=sha256:cf5468a02ee221490fadf43108ef663d74a63efddf9ae5102748511c568f0b82

Observation 6e71dfec-1ae2-40e0-934d-9d0c8eee30b4 · outbound

This paper cites and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Sun, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.955298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.884469Z digest=sha256:46c77748d7637f674cf1cb327fa6e21b9b057c2533709eed2f2c3088112518b9

Observation 92e305af-7512-4a3b-85d2-659758cb44ae · outbound

This paper cites Principled preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Principled preferential bayesian optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.700299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.981051Z digest=sha256:5f93a02638b38feb7091795ce16c86a84f68ed97072d317ce59663835b47f3a1

Observation bb3d5c7e-ab48-4452-b719-d08bd5b063cb · outbound

This paper cites Zeroth order non-convex optimization with dueling-choice bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Zeroth order non-convex optimization with dueling-choice bandits

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.449181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.103390Z digest=sha256:30b33d4f01ac2dc4661dfe79ffb03c566bcadcf2f93f2aa553ca686e0cdfb0f3

Observation f840f36d-5b10-4be3-8f71-ee7fb1ae6126 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based reinforcement learning with finite-time guarantees

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.176236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.192620Z digest=sha256:8db5b09abe6987a721f77cd40ca48d5c5716256c16f779cb485cacfa17e0505a

Observation 9f09a41f-a9d6-47d5-91eb-6c09adb7e6d7 · outbound

This paper cites and Joachims, T.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Joachims, T

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.923255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.253187Z digest=sha256:96b10af848cc7a75720aeaac84254008f2948e99e09ea6e45b02e6b9d91b2631

Observation f90a452d-7f1b-4882-ade5-a0cca8fe9794 · outbound

This paper cites The k-armed dueling bandits problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds The k-armed dueling bandits problem

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.669843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.318578Z digest=sha256:7504a8630826b69cff2ad8be6e0dc0752f925b28581c84f21a1ad8a3166bc8d1

Observation 084eb858-f460-4be7-b7e6-405729726a10 · outbound

This paper cites D., and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds D., and Sun, W

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.427524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.418047Z digest=sha256:472421ac9a7bfabd1ffd0269b008eea9b1fc8ff76d16a7a641b2c207e92f3a17

Observation 9d3dc523-5eee-4577-b38a-cb2a093b2b88 · outbound

This paper cites Relative upper confidence bound for the k-armed dueling bandit problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Relative upper confidence bound for the k-armed dueling bandit problem

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.179608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.531566Z digest=sha256:1388a8b943cdea6575490910fbd953dd115e7032aa08190acaf06a6b3380b05f

Observation 243d5c50-0041-43b5-9901-9c4cb9c5cdb0 · outbound

This paper cites Mergerucb: A method for large-scale online ranker evaluation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mergerucb: A method for large-scale online ranker evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:45.950061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.634287Z digest=sha256:4e3e0abb96010931d78051fe8a752ffe61a428e37a7d1438df7b7b8a089efe13

Pith citing papers

No inbound Pith citation observations are available.