Pith. sign in

Paper Citation Record · LEDGER

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds

As of 10 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23673 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:45.634287Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbb2eea2-eefa-4813-a6ce-abe16523ff2b · outbound

This paper cites write newline.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:39.878545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:39.878545Z digest=sha256:b1999046004bf819740a06d868c627179b45d9eaba20cdc0a2599b1217c86695

Observation f25d9bbb-539f-4ffa-a943-5380cdd1e2ee · outbound

This paper cites Online learning for linearly parametrized control problems.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Online learning for linearly parametrized control problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:57.055986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:39.976257Z digest=sha256:4eea913ff197704e6b20111c25d88cde79f75a55fecfaea6f83dedfee278f6d3

Observation f63f9162-3596-4bc2-b4e7-b14c5fb6ed7d · outbound

This paper cites Reducing dueling bandits to cardinal bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Reducing dueling bandits to cardinal bandits

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.865663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.043920Z digest=sha256:e3a7d85a144d0a1fc5cb0aced5395106d7638dbdcabea88333bcb381bcb8b9b0

Observation c916829b-9808-4ebb-a0ce-b4bda5cf7a20 · outbound

This paper cites S., Hu, W., Li, Z., Salakhutdinov, R.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds S., Hu, W., Li, Z., Salakhutdinov, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.138320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.138320Z digest=sha256:cff83c6700c6f3140b3d90228eeaa6a5a8e6a6c9a34994a19b1e2a2dd9b9256c

Observation 661e7698-f079-457d-91a0-dcbd759a411c · outbound

This paper cites R., Daulton, S., Letham, B., Wilson, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Daulton, S., Letham, B., Wilson, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.750917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.241854Z digest=sha256:34dc951c4516c0b77baf18b81243c6b60876ef49aa6909f12a2b13518f2ed4cc

Observation 3c8de591-d8e2-458b-9040-58d0cb36bd85 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based online learning with dueling bandits: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.449269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.361927Z digest=sha256:3896634d820d83d51ba6a48f85a71bffc5081c2f733d45dee289b88632234010

Observation 3b08b2bb-24a7-42e8-af6a-df979b24cc66 · outbound

This paper cites Stochastic contextual dueling bandits under linear stochastic transitivity models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Stochastic contextual dueling bandits under linear stochastic transitivity models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.187735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.442785Z digest=sha256:681b70b3436b4bb10d3e4ac9a36bb7c851649db9952fbf9ad9227d995190391c

Observation c5f9d829-ad58-45fd-b147-a251fe5e750a · outbound

This paper cites Mat \'e rn gaussian processes on riemannian manifolds.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mat \'e rn gaussian processes on riemannian manifolds

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.922717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.521123Z digest=sha256:3d35b44331a03bf2a52f013a75498cf7d9d9507a1607a6b32df64588197871e6

Observation 1cbd7aa1-5dfe-4a1d-9f66-859d3c207769 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.628781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.628781Z digest=sha256:68570988f1fc2c61399f994ade7bdba04c47e59905205e07828bf3a8d586745d

Observation d1b04aac-25ff-468c-a4f1-279ce8a31e28 · outbound

This paper cites A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.734979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.734979Z digest=sha256:578ec8228f338584eee3e7e553b749b6775c7d9b1b3f1e58299a95c248ff3663

Observation eb9496c5-f672-4075-a016-e37658ba9289 · outbound

This paper cites Instructzero: Efficient instruction optimization for black-box large language models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Instructzero: Efficient instruction optimization for black-box large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.729489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.835746Z digest=sha256:3a1cba88c5b9bebe16484143aa4ada74d36650664634cea893f54aa395361f79

Observation 72f068b0-e951-4dec-aabc-ef40b7e314d5 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.901099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.901099Z digest=sha256:be0227e60ac282e40937c5dcf7181d5871de551e1412298e6aad09beedb93142

Observation cd4f5627-73e6-42a9-9a40-4be5c0d094e4 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:55.354090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.994901Z digest=sha256:8c9c19603339d70708dd5c569e85cf3120fc3bba6f104f14271e18f7cbc95f47

Observation 3a2067e6-4446-4226-8812-ca1d198dd0cc · outbound

This paper cites and Steinwart, I.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Steinwart, I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.009181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.069695Z digest=sha256:57ee40ee5d143fecd335a2bd796fae6947fd54735e23a240e9913d0f2a4acb69

Observation bf1ae0a8-8f2e-4f09-9963-fb340e6a3851 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.743049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.168084Z digest=sha256:0cb578388f26db0de187f4a003189b762bc54e0ffd54cf5bcdc0f84ceb5ed947

Observation 6dd47c75-eda3-44a6-93c1-9020fbd86d56 · outbound

This paper cites E., Slivkins, A., and Zoghi, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds E., Slivkins, A., and Zoghi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:54.433457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.269844Z digest=sha256:f89b0fa92e4706276dfee0e3df9aa799e79c201b4f58621d198b560292016340

Observation 279f94d4-b194-4864-b290-d74173a76ca7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.070394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.351047Z digest=sha256:91f53149fb7a402e4a79581cd0c8810d58637e9e2df52a161280346277c23162

Observation 90cc3a77-2405-4f42-bf4c-0d4f49d85df7 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Improved optimistic algorithms for logistic bandits

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.772960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.453532Z digest=sha256:b481bb8793de0072a6a7152796710f370edfaa17463f38ab5ef8aa0c25dc2ede

Observation 3fd1dca9-b811-4db7-8ba6-edaa73054cf9 · outbound

This paper cites A Tutorial on Bayesian Optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:41.544682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:41.544682Z digest=sha256:9adfdefaaece83817967f6b2fdfcd58b8416ea5ac1e1f937ca83de88f2338229

Observation 18fce1e8-13d2-40ae-9586-1d9ec0b24e46 · outbound

This paper cites R., Pleiss, G., Bindel, D., Weinberger, K.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Pleiss, G., Bindel, D., Weinberger, K

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.480243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.638418Z digest=sha256:b8737419461cb9879c6dbe2a0892a11263ea2b840476f5a10cc0339b95b7fd85

Observation b5164a3c-46a1-4518-9082-846da787c71c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:53.193134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.704728Z digest=sha256:a94320e79f52106b320681a1be86e9360eefcad3c13de19cf6c1cb53e14c03ae

Observation 1cc4fddb-d77a-422f-bbb7-62bba9674992 · outbound

This paper cites L., and Thomaz, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds L., and Thomaz, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.899224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.773886Z digest=sha256:38a2249c45c301c079f08d8a03902b5089570e8ff59c98e64418e2786aa9dd66

Observation f82de9f4-ae72-47bb-a88c-9b04475e04bf · outbound

This paper cites and Yang, X.-S.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Yang, X.-S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.600042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.930302Z digest=sha256:1a08724357921f692f01282f39d82f43bc17239bd51b3912de6c91ecb06ad23a

Observation dad41594-2365-4926-9a1e-5a488c4bb929 · outbound

This paper cites R., Schonlau, M., and Welch, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Schonlau, M., and Welch, W

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.054479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.054479Z digest=sha256:03f45b7e245dac12df289ad43bcf7f658dab2d7cf118c0e0914dce45ccf4f595

Observation 807202dd-e4de-4525-b1ca-f3f9eeb24320 · outbound

This paper cites Feel-good thompson sampling for contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Feel-good thompson sampling for contextual dueling bandits

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.335804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.145030Z digest=sha256:552c7bfee0cc257b61d4a19068a6e9ca77210bec4a1c6c364091eff5e1b6154f

Observation 70eed8b9-bb8e-4b6d-b810-cfc21206ee71 · outbound

This paper cites and Scarlett, J.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Scarlett, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.109171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.274206Z digest=sha256:f0c9011de94be7d9f6828e7f2fd32d32fee3f196fa1bbe2df573593fd198cd3e

Observation 6fa7c559-fe37-473b-916b-3147b64858b7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.870471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.368045Z digest=sha256:821b44c6c5dc570d6f7b5aeddf386c2734d0b2c2f002e06f94e50550d3f07beb

Observation c52c829c-9edc-4326-9bd9-17ec706ccd74 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.603727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.497274Z digest=sha256:dead3c32ad41f9d3e38e193c8ce6d162179f82ba54e6e5fdd593c4cd70103271

Observation 3e05f0f4-440a-4c96-815a-b223275df31c · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.595875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.595875Z digest=sha256:44245a61641f63456f7f4ed2a6982d751423f6d409e4cb98d02d8053f53fa1e6

Observation 05359119-0546-450b-85ca-779f84c2d013 · outbound

This paper cites Kernelized offline contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Kernelized offline contextual dueling bandits

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.302536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.722443Z digest=sha256:0960ef0b92d60e55e6ed2675bbbc4b9f568c3101a8a48efaab743f38dcf8cf13

Observation 5a862fc4-f9e7-4c39-a3ff-b6c87c6ebf2d · outbound

This paper cites Functions of positive and negative type, and their connection with the theory of integral equations.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Functions of positive and negative type, and their connection with the theory of integral equations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.069454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.844443Z digest=sha256:ad32053d8545a36446039ffb3cfec72ec23bda18e255ed29ab128bd031eee0d3

Observation 1ce9876e-4bd5-4bb8-86be-301ac54aa385 · outbound

This paper cites Projective preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Projective preferential bayesian optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.884143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.936733Z digest=sha256:5de12908f63749bba624d0d88301c99870b3db5a0774701040e99c24a9a4f5fa

Observation 849b6473-3f9d-4ee8-883c-119846cf8701 · outbound

This paper cites Dueling posterior sampling for preference-based reinforcement learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling posterior sampling for preference-based reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.040733Z digest=sha256:af6caa9ddf7a22ddd13d4facc2e2bab5ad33c065d38c6a49f65e509f6c660737

Observation 8646b46f-3fbe-4102-89a5-ea926abb3064 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.162585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.162585Z digest=sha256:4e57054d2b8889763b07b5d37657b4e1de9afb8b12c578632e910853cc28cc21

Observation 4f7d64a3-69ac-49b3-a6b5-ab2d59de08c1 · outbound

This paper cites Bandits with preference feedback: A stackelberg game perspective.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Bandits with preference feedback: A stackelberg game perspective

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.681901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.268258Z digest=sha256:3b8add7b18be035dcdf4d23c6956dea394d9c119ad731270ae9ee642c7f71233

Observation caa9fe46-d605-42c5-bc2a-bcebf3c6eac5 · outbound

This paper cites Scikit-learn: Machine learning in P ython.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Scikit-learn: Machine learning in P ython

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.415191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.415191Z digest=sha256:f4568ea770af02e0d2e97b9bddb90bf09d04501a677f1fca0e1dc9bef7da5055

Observation 166d52dd-32ae-4492-b81d-5df8ea1d1138 · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal algorithms for stochastic contextual preference bandits

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.504449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.528257Z digest=sha256:98193a073dddd8129291582d68669bf8356c05ea3ffb13bb6e94b65add3b67e1

Observation 7d52213e-1a68-489b-9e55-ee72fe4bb353 · outbound

This paper cites and Krishnamurthy, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Krishnamurthy, A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.376160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.638184Z digest=sha256:72eb9546e560e0698985ad0f5107a4434af9faceb1b0e1ca54b64e587c09161e

Observation 889bb251-af6e-403a-9647-950bd1ee4e92 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling rl: Reinforcement learning with trajectory preferences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.226512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.712472Z digest=sha256:785bd9efed1ab325e3b46437baabbf981625405d17772b79e34aa17e8d1d8fdd

Observation 1c144d77-19f3-46d8-b247-76c1035bb965 · outbound

This paper cites A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.071549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.782700Z digest=sha256:adea6238545715a79e46172de0d9bc597ebf3205ab069b3a5ff13b60c54e1318

Observation c2c9467f-1cca-4af2-bb75-07d3faab1062 · outbound

This paper cites Lower bounds on regret for noisy gaussian process bandit optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Lower bounds on regret for noisy gaussian process bandit optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.875455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.853547Z digest=sha256:f1abf20e30c32e924b3a31cb62a4762cc85f78d7dbcca94a50339b7ac1cf6953

Observation b123e8fe-ebfc-4bcf-940c-41a2afa6489c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.701942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.935161Z digest=sha256:1c56124168437663ab1393e603f0e3791d4eaabd24752ffebcd787bb7a1fbdd7

Observation 3758a7e0-2e92-465b-bbf7-cdeb81572b1c · outbound

This paper cites P., and De Freitas, N.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds P., and De Freitas, N

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.036902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.036902Z digest=sha256:5a7ba6e930084b757fb3c9c00735c20228e87fbd0d876b8eb4960c9f92445e1c

Observation 23b44e2a-b881-4f5f-981d-8fbcc8ea02c5 · outbound

This paper cites M., and Seeger, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds M., and Seeger, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.499940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.117889Z digest=sha256:bbccf2e06be1470537948bb1cd6be5bcfafcc46af89305302c5f823b31e25f9b

Observation e8d65b47-20a1-4c62-9db7-57bebf343634 · outbound

This paper cites Towards practical preferential bayesian optimization with skew gaussian processes.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Towards practical preferential bayesian optimization with skew gaussian processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.311254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.202090Z digest=sha256:46d39ffe58fc3176f2d0b01e42d91509009fb8aa6ea200be0bafee513e6e8c55

Observation a985fd11-9b1a-4b58-a2eb-372ec3ca4b69 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.159137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.270455Z digest=sha256:ae2276573f465c8a6c616591b92364bb171f8633e7290946a2758f3e3535b3b0

Observation 08434648-ed7d-40f3-bf5d-6b6810b15521 · outbound

This paper cites Optimal order simple regret for gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal order simple regret for gaussian process bandits

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.991095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.350553Z digest=sha256:f1b8c06697f76a95e7ccafdd51383e6923a18d035978ebe3a172794b67f5c4a0

Observation b9381b72-7c92-4d46-9a8c-8c9188bc71d4 · outbound

This paper cites On information gain and regret bounds in gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds On information gain and regret bounds in gaussian process bandits

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.847219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.435530Z digest=sha256:03d6579305962e2834d2c0465ce77a78955b5648e6418336aade4f223aaafc7e

Observation 4d154dbc-42cf-4bea-8214-ebb4695eacbe · outbound

This paper cites Finite-time analysis of kernelised contextual bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Finite-time analysis of kernelised contextual bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.739454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.531719Z digest=sha256:398ce24215badafb41f749a14f20fd1b57ddcd9de00dea31a3d34c6202c559f7

Observation b91b39b6-17ce-4ffc-a485-0509836099e9 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.477948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.601543Z digest=sha256:dce490b3adcd298acbcb0acef9ce4fa2b7708ecf7088708346855e4e0c55285e

Observation 4fd5c7cf-d33e-4877-9957-135aea590e2a · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds High-dimensional probability: An introduction with applications in data science, volume 47

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.680747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.680747Z digest=sha256:c3668f0f358bde49d3910ae936863c61e1df3bd123b2794c5fd5398a11c00dfc

Observation 25dd1c45-fc65-4544-b0e1-c48f8c1ee64d · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.210392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.789087Z digest=sha256:89af6eac7b18da8e493df0f29d30662a181e1d7cd633993fb47d37a9a610241b

Observation 6e71dfec-1ae2-40e0-934d-9d0c8eee30b4 · outbound

This paper cites and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Sun, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.955298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.884469Z digest=sha256:aa229073b7fa27933cff80f6c8f1ebf9a42b80d43a2d29db2ea6a2c5e22d4714

Observation 92e305af-7512-4a3b-85d2-659758cb44ae · outbound

This paper cites Principled preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Principled preferential bayesian optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.700299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.981051Z digest=sha256:71bf0edb2674fe79a6fac55e4d1cbc62288ad0bc81efc3d28956c188c0912473

Observation bb3d5c7e-ab48-4452-b719-d08bd5b063cb · outbound

This paper cites Zeroth order non-convex optimization with dueling-choice bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Zeroth order non-convex optimization with dueling-choice bandits

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.449181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.103390Z digest=sha256:6d01ec4f9a77582d463bf0e624d87d9d9bda2632530e0ab544e3dca4e451e737

Observation f840f36d-5b10-4be3-8f71-ee7fb1ae6126 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based reinforcement learning with finite-time guarantees

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.176236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.192620Z digest=sha256:3c4ed45d68d5335e5af1962ecd9c3a9b2890af60ba2632d1742cb782c4298969

Observation 9f09a41f-a9d6-47d5-91eb-6c09adb7e6d7 · outbound

This paper cites and Joachims, T.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Joachims, T

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.923255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.253187Z digest=sha256:867ffb72f113e0f2875e0c3b9973057739cdd9819f46212bcb100ae376d2acd2

Observation f90a452d-7f1b-4882-ade5-a0cca8fe9794 · outbound

This paper cites The k-armed dueling bandits problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds The k-armed dueling bandits problem

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.669843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.318578Z digest=sha256:06f6f15026a24be2ac78392df7ec7cd2f9062bf2a8777a4929209fc9b1f2b9ce

Observation 084eb858-f460-4be7-b7e6-405729726a10 · outbound

This paper cites D., and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds D., and Sun, W

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.427524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.418047Z digest=sha256:e3300b45f4f330bbe44610a74293739f8ca78082cfbcb2e4ed8c30cf51b3101d

Observation 9d3dc523-5eee-4577-b38a-cb2a093b2b88 · outbound

This paper cites Relative upper confidence bound for the k-armed dueling bandit problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Relative upper confidence bound for the k-armed dueling bandit problem

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.179608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.531566Z digest=sha256:7159a0df5f5e8fb5e96d9b1300a371c902ad97cde9aab6d5e61832123de1a6aa

Observation 243d5c50-0041-43b5-9901-9c4cb9c5cdb0 · outbound

This paper cites Mergerucb: A method for large-scale online ranker evaluation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mergerucb: A method for large-scale online ranker evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:45.950061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.634287Z digest=sha256:289e759b90dcdbefff8cb282b0a41552c1a6c67bc8ff78f9b2d95c46349dae06

Pith citing papers

No inbound Pith citation observations are available.