Pith. sign in

Paper Citation Record · LEDGER

Combinatorial Reinforcement Learning with Preference Feedback

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.10158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10158 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:57:04.381612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T19:57:04.556666Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b18648c4-8d23-482a-b0d4-776f4aea6e69 · outbound

This paper cites Instance-wise minimax-optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Instance-wise minimax-optimal algorithms for logistic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.791208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.224384Z digest=sha256:9be707f02848d94148366b75582491c9a8929b9e5c1edff9555faf3e4ab56fa7

Observation 28165428-c5d9-4984-b1fe-ef7bca76f6de · outbound

This paper cites Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.761100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.231128Z digest=sha256:589c5f8da6a1fc5e01863767a71cd67cb0fb54398ea0cc77b599eff42019a75d

Observation 2f556e53-d9b4-478a-93e5-6bad9c169184 · outbound

This paper cites A tractable online learning algorithm for the multinomial logit contextual bandit.

Combinatorial Reinforcement Learning with Preference Feedback A tractable online learning algorithm for the multinomial logit contextual bandit

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.731998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.237577Z digest=sha256:858f98e07f32606eac179d53214d1941fd331dd45df8a3a0877982285010ff14

Observation 548ffc32-b4db-40e6-afa2-84473a54d7ee · outbound

This paper cites and Goyal, N.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, N

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.706758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.245239Z digest=sha256:b3ea8a55def0238941af1ac54966a887a1b55e0d324ddf95c424b74dadfa8179

Observation c5559935-c99c-4704-94fc-f6e73a4e3652 · outbound

This paper cites Thompson sampling for the mnl-bandit.

Combinatorial Reinforcement Learning with Preference Feedback Thompson sampling for the mnl-bandit

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.681536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.256779Z digest=sha256:376f0ded1fa4d3c41063706eacf83fb138d550cc31674bd3967d7549c69fafd4

Observation 0080f2ad-5033-432b-92f4-b3169a7009ab · outbound

This paper cites Mnl-bandit: A dynamic learning approach to assortment selection.

Combinatorial Reinforcement Learning with Preference Feedback Mnl-bandit: A dynamic learning approach to assortment selection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.262831Z digest=sha256:84f39af5462886f508b8eee353641977fdcd88a8f36c44d581e9a7db2ff0905f

Observation 1a6b6990-e2f1-4953-b1bb-3f9504f68356 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Combinatorial Reinforcement Learning with Preference Feedback April: Active preference learning-based reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.271517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.271517Z digest=sha256:ee8e24a15c307ff3f6fd05569f6c44e4076425d918c2bd836e5b891f01a0b425

Observation 6ce715ec-61b0-468f-9f9a-9d88960490eb · outbound

This paper cites and Thrampoulidis, C.

Combinatorial Reinforcement Learning with Preference Feedback and Thrampoulidis, C

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.611297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.284760Z digest=sha256:c096782e64429f0c9fc2153d0d34aae3548bbb1a6529ce83702a4016205618e4

Observation 87e0a6b4-94a7-4b9e-af07-8a22e8837bc7 · outbound

This paper cites Distributional off-policy evaluation for slate recommendations.

Combinatorial Reinforcement Learning with Preference Feedback Distributional off-policy evaluation for slate recommendations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.587861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.292863Z digest=sha256:48b201b1a983f8a65d3cbd16cf0fe11650c42c6cbfdfdae9dec3796bbbe7be5f

Observation 280622c9-3417-4960-854e-bad58b1406bd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.561019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.304989Z digest=sha256:dcb9f49a70803705fc14d87487221400c4bfbc0ec96eb45e5176a4faebcde764

Observation 2ba47c7e-b279-4d54-8431-699597d2ea37 · outbound

This paper cites Combinatorial multi-armed bandit: General framework and applications.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial multi-armed bandit: General framework and applications

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.532539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.312614Z digest=sha256:a5d3f3a6fc7fdd7ba44816237996c904a219c32e0252ca3e54c75a49d1972be6

Observation f373dd7e-7ebb-408d-b664-499d8c7959c3 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.507919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.323399Z digest=sha256:6c450aa4397df84aa3a91d6b1ad8a569a6839c6834142e1a1347f336503aa916

Observation 1301a24b-9616-49e4-be44-00e3b0eebef5 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Combinatorial Reinforcement Learning with Preference Feedback F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.332237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.332237Z digest=sha256:23cf4dd312ef34a5738bb8a111c4e767963118c68704d9a1bfef4a8393b0f7ea

Observation 9f827c34-880e-4eb0-b828-1d6e72e67e0d · outbound

This paper cites S., Proutiere, A., et al.

Combinatorial Reinforcement Learning with Preference Feedback S., Proutiere, A., et al

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.443014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.341524Z digest=sha256:940d07e81d0b52b4d43dd71942ce8453f2db52c94574c1123ec6b3bedef37d60

Observation f486ebc0-5827-49a4-b03e-e77ade3febf1 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.352173Z digest=sha256:164678e3d12d645f43e1eb07a34cc9f2050763a90016d15a006f611a254b71b1

Observation 6efd8f5d-8ecd-4f79-a205-07b39d9a6b00 · outbound

This paper cites Assortment planning under the multinomial logit model with totally unimodular constraint structures.

Combinatorial Reinforcement Learning with Preference Feedback Assortment planning under the multinomial logit model with totally unimodular constraint structures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.383925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.359505Z digest=sha256:bbad9bd741aec245a078e2c5dc3bc00c07138924b207e288743e8aa415a4a3b7

Observation d28026fd-a4ba-4d6e-99d7-21851d4eb11f · outbound

This paper cites Reinforcement learning with combinatorial actions: An application to vehicle routing.

Combinatorial Reinforcement Learning with Preference Feedback Reinforcement learning with combinatorial actions: An application to vehicle routing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.337692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.365124Z digest=sha256:d37ec64d150a17b145870a1770a2ae688febb9fa5007670eeb40b3d7ea7ebca8

Observation dfac6861-d89d-44c1-9e18-c0cfc3117ede · outbound

This paper cites Bilinear classes: A structural framework for provable generalization in rl.

Combinatorial Reinforcement Learning with Preference Feedback Bilinear classes: A structural framework for provable generalization in rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.314252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.371484Z digest=sha256:ac1aa4022b98b649c7febe06638f45505d8340e83394903b69e20aa587a32514

Observation ac33eec9-1542-41d9-aa9a-81687450b21a · outbound

This paper cites Cascading Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Cascading Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.850736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.377776Z digest=sha256:71ed0e051944babc6b4a512eea66becd134a9f0d0c0873ea35d52b8162a20367

Observation cfa91a6f-ffe0-4abc-946c-00885a286f88 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Improved optimistic algorithms for logistic bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.405443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.405443Z digest=sha256:c4d81fb938c7d4b8b893e671548c8caed10bb7f9306c5471863f37b3e743c292

Observation 098c220a-6abd-4050-b907-6c19d2a35cb0 · outbound

This paper cites Jointly efficient and optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Jointly efficient and optimal algorithms for logistic bandits

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.269061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.475614Z digest=sha256:73bab628b58ae498bce4d7a8d9fe3f77a6bedc342c4fae749660033c36c5fcd3

Observation 7e340fbd-91da-4e35-ba4b-b239c9866093 · outbound

This paper cites Parametric bandits: The generalized linear case.

Combinatorial Reinforcement Learning with Preference Feedback Parametric bandits: The generalized linear case

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.243063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.546882Z digest=sha256:fe456af4675ff819a0987cd65eeb7c881d1ea2b984f2d09ed24fcf075cf7ca52

Observation 2f8a65c8-fb8b-4f3f-8849-0f9032f82e28 · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Combinatorial Reinforcement Learning with Preference Feedback The Statistical Complexity of Interactive Decision Making

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.598167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.598167Z digest=sha256:a7ab9feae59fc645897c55940c6b78443a8be432a78da0bd6a406af784e51fd4

Observation c03b3cd5-d973-41bd-a51c-8dfa139917cd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.218068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.655974Z digest=sha256:2e4b74ac1c69b59db786582764b9f9b2431389b363caa9136549d607b41299e5

Observation 6928cb0d-781e-413a-8e6e-441d96001d7d · outbound

This paper cites Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.748595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.665758Z digest=sha256:fc6a3bd07ec7bc1f218517602c4b10fd67eadc994f08f43f48e31dae6c7fd891

Observation 3d67be2f-445b-4ee9-8b72-fb62b692493b · outbound

This paper cites Slateq: A tractable decomposition for reinforcement learning with recommendation sets.

Combinatorial Reinforcement Learning with Preference Feedback Slateq: A tractable decomposition for reinforcement learning with recommendation sets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.197624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.672769Z digest=sha256:3661420daf68df938a7bda69507f1d02b2d10d1f79326a971e480070b3016dfb

Observation db054e73-4680-4d27-87cf-20070c8075d5 · outbound

This paper cites Randomized exploration in reinforcement learning with general value function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Randomized exploration in reinforcement learning with general value function approximation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.680375Z digest=sha256:0ac8fe89c26718932e8fe9347f93ecb05202d4c2156f3033919803bd2bb9c6a0

Observation d37ac1b4-fda5-476f-a172-b224f7503961 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.133885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.685870Z digest=sha256:28e840a916d7767634a919d3de7d52a67a9adf3a8a1a07849ccfda10950de4ae

Observation eb345469-8408-4da7-ab13-a4e73fec3f43 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.102859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.692642Z digest=sha256:923b28d551f881d4121a207f27f175b2d986c9b620572431d32f7bda355bd731

Observation 768a359f-e859-454d-8b5d-64a5cc65c0f9 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.076408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.699888Z digest=sha256:d170c010f3fb05052db4996caed62f6396f824e70e7889ebdce6604298739358

Observation 20505fb5-e7b5-4bb6-bba8-f729ebe59dd9 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Combinatorial Reinforcement Learning with Preference Feedback Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.053824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.705775Z digest=sha256:10ed8b32ec3163312ec1a71044c4ef2d5c53da637a6f68429814ffe5c8b40fac

Observation c7559e90-d365-4d5d-887e-5c33a19d1edf · outbound

This paper cites Online Sub-Sampling for Reinforcement Learning with General Function Approximation.

Combinatorial Reinforcement Learning with Preference Feedback Online Sub-Sampling for Reinforcement Learning with General Function Approximation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.711779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.711779Z digest=sha256:5e7c30da01d464522e919161019f7b5186388b4a5dcf11efe358111e89b17444

Observation e51f6d31-7122-487c-b078-3f5864f46902 · outbound

This paper cites Cascading bandits: Learning to rank in the cascade model.

Combinatorial Reinforcement Learning with Preference Feedback Cascading bandits: Learning to rank in the cascade model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.022919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.717594Z digest=sha256:900c8f31ac89c050350c36d358436605d6545f5de45836af22ea28f370c767e7

Observation 97c5c92f-6b79-4add-a98d-4cb47ffe6888 · outbound

This paper cites Combinatorial cascading bandits.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial cascading bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.990933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.722642Z digest=sha256:834cd799d395554001b3f846684529f13018cc912fc501c3b339593a925cab89

Observation a616e48a-f068-49d1-b623-93d9fbe07a2e · outbound

This paper cites and Hutter, M.

Combinatorial Reinforcement Learning with Preference Feedback and Hutter, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.969336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.727018Z digest=sha256:201f01d320a4aded198c17da859527960d2990e6ca66d17ec140cbfb0274ba12

Observation 1d65c601-5eb0-47a9-a4f8-b27ae658b390 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.946498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.732283Z digest=sha256:125f798a92d60f8442fb3ef1c7ddbb51f793076ad46ff30b30d193a462391878

Observation 28b636af-c5f1-4f43-896a-d20548e80e13 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.736731Z digest=sha256:d21e16c561f49a3bdf5dce7a8732476730043ec55db91b54b796d86e20eb24ec

Observation aabc5773-222e-482b-97a6-cc8434b97888 · outbound

This paper cites Online learning to rank with features.

Combinatorial Reinforcement Learning with Preference Feedback Online learning to rank with features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.741426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.741426Z digest=sha256:2d062f886fcdbc6180f0e89859465aee963e268a07d8476e7e4aa0b39cb0f582

Observation 2e7da764-f091-4cd8-8af8-a1b6b9e12d03 · outbound

This paper cites Modelling the choice of residential location.

Combinatorial Reinforcement Learning with Preference Feedback Modelling the choice of residential location

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.869998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.747449Z digest=sha256:1200476ce8d8b0c40a64d00d7cf2f1822be94b85c75c948878652ef109866cce

Observation 7ea4aeb3-84a7-439d-8dd2-8f0fb9d3985f · outbound

This paper cites Counterfactual evaluation of slate recommendations with sequential reward interactions.

Combinatorial Reinforcement Learning with Preference Feedback Counterfactual evaluation of slate recommendations with sequential reward interactions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.846788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.754104Z digest=sha256:5b6f8c5b28d795b620243b5c9dd0958fa152439fa8a4ce856a4f313d08149ec5

Observation fcfff1cd-f4ab-4158-ae4a-b06c3f79319c · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Combinatorial Reinforcement Learning with Preference Feedback Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.761945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.761945Z digest=sha256:85a6fb32955dff08076e604bdb95b4f3d5a705bc5861728b019c18a1aa091db4

Observation f28ccee3-682d-4983-af4c-fc167647d0c5 · outbound

This paper cites M., and Van Erven, T.

Combinatorial Reinforcement Learning with Preference Feedback M., and Van Erven, T

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.823133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.792670Z digest=sha256:4d4ecafbd97f7eabe242b5397b66bc1524b2ff8f544b4cd8fc68d56b9a1f9176

Observation 56c2cf88-add8-4503-afae-f42231ec6f72 · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.798459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.825505Z digest=sha256:72c8d25ca08ce287b01c0b9303220efe8b350c82c7876963629c80eb68025598

Observation 817c122c-11e5-4d5e-a8d0-f173f9c3e92f · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.776371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.844502Z digest=sha256:f5289bb9a43af457cf3467f4490acbe19a27709a82463a7d81be32d1a084be50

Observation 5709ed03-ef9c-4ebf-8579-29b622ec9063 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Combinatorial Reinforcement Learning with Preference Feedback Online Learning: A Modern Introduction Using Convex Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.870065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.870065Z digest=sha256:ac96341418294338ffb3ed27ef2fb8da3eedc811347bdef12743ab7a17d0002e

Observation 71557d63-5f85-40df-aeba-389f1c2d1a42 · outbound

This paper cites Training language models to follow instructions with human feedback.

Combinatorial Reinforcement Learning with Preference Feedback Training language models to follow instructions with human feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.896163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.896163Z digest=sha256:87cd75f35e80a326ea14068b2cfc781c01fef2b1cbe7bc710d390b86309783cc

Observation b44bb0a4-8212-49ba-821f-f13f5e3bc851 · outbound

This paper cites and Goyal, V.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.725451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.904201Z digest=sha256:b657fb3b28e8a8a4f05f9d7d556ce94851b202418e74ea08a4593a5b23654971

Observation a975f62f-39bd-43ad-aa70-cfa62807d193 · outbound

This paper cites M., and Shmoys, D.

Combinatorial Reinforcement Learning with Preference Feedback M., and Shmoys, D

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.702565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.910945Z digest=sha256:27d9cf39f327c6accdb8803f78454cd0b0f56af7071ae8abb76f18918adb279e

Observation fd73fb0f-6618-47e3-8b24-2fa8b5fc3fcd · outbound

This paper cites and Van Roy, B.

Combinatorial Reinforcement Learning with Preference Feedback and Van Roy, B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.674460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.918859Z digest=sha256:ae033e263c08bc29945c7a78a2adc8863d144314f05cdb98a78153508ea8af7a

Observation 3b9cb725-dbc2-4f30-886e-b86fde07259b · outbound

This paper cites CAQL: Continuous Action Q-Learning.

Combinatorial Reinforcement Learning with Preference Feedback CAQL: Continuous Action Q-Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.926888Z digest=sha256:744bd923e43d23212bb01b7ae95815f13facefbc1984e16f060f2eb636a1a9a1

Observation c361c8e7-8103-4598-9c9c-54a000858ecd · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Combinatorial Reinforcement Learning with Preference Feedback Dueling rl: Reinforcement learning with trajectory preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.934451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.934451Z digest=sha256:b0dccd2b1fb4f0a5e9b0a0861b88eae9a1269b4506e2ed2cb30685e92f226f6c

Observation 0f14d212-1104-4eb3-a749-0100b3ac460c · outbound

This paper cites and Zeevi, A.

Combinatorial Reinforcement Learning with Preference Feedback and Zeevi, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.571717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.943188Z digest=sha256:be0af30692eb584e7a9fe94e5874eb54d4b69ddfd8aebd4fa845777da18cd71e

Observation e54c55dc-8a4a-4da0-b233-80cb91fa2068 · outbound

This paper cites Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.951927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.951927Z digest=sha256:7c28e8c2a4fbd7b8705b78882d10a43bc3a9dc207cc5a54f21c0a03b67d59b21

Observation 2d6c0153-9fd0-48af-aeb0-753d3caa5d61 · outbound

This paper cites Off-policy evaluation for slate recommendation.

Combinatorial Reinforcement Learning with Preference Feedback Off-policy evaluation for slate recommendation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.490085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.958603Z digest=sha256:656aab186ebd22f532508411f96c20bc382f7951baccf94627f93cb2a8c7c1e2

Observation 966e96de-4d75-4ed4-a76e-7b8f27233eb2 · outbound

This paper cites Composite convex minimization involving self-concordant-like cost functions.

Combinatorial Reinforcement Learning with Preference Feedback Composite convex minimization involving self-concordant-like cost functions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.392696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.967403Z digest=sha256:6e84fb5f6ba2d965151d4f560af01299313f2e59d526b9a8f1c963dfe0e50457

Observation 0381e4e0-bdaa-41eb-8fd2-05ade2ab0bc1 · outbound

This paper cites Control variates for slate off-policy evaluation.

Combinatorial Reinforcement Learning with Preference Feedback Control variates for slate off-policy evaluation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.370807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.973733Z digest=sha256:3e5fe60a65cee4107b826911fd5df035ad68c3145f4f38343087ca0444873821

Observation 94a32340-6a5b-4085-bb4e-f7ca283eb595 · outbound

This paper cites R., and Yang, L.

Combinatorial Reinforcement Learning with Preference Feedback R., and Yang, L

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.343071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.979355Z digest=sha256:ffa293a0b8fe01df6d8c256787b9bcb53648eccd035249927bf21d5a99687f51

Observation cf765a08-75ae-4f55-8dba-ffcf813dc963 · outbound

This paper cites S., and Krishnamurthy, A.

Combinatorial Reinforcement Learning with Preference Feedback S., and Krishnamurthy, A

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.299861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.984567Z digest=sha256:763d2f4d3041a628e2ee7cc16944701121da8607591e8080e7400d4e53a9f007

Observation f7cfa02b-79b3-4be3-9004-935ad45b2f17 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Combinatorial Reinforcement Learning with Preference Feedback A survey of preference-based reinforcement learning methods

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.992315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.992315Z digest=sha256:f0a9a2f4faf987fac6ffd8c03852057253404b6b2197d103b14da222343d9d72

Observation 709c594c-281b-4b1f-b016-6bb1a69643b8 · outbound

This paper cites and Wang, M.

Combinatorial Reinforcement Learning with Preference Feedback and Wang, M

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.250244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.998266Z digest=sha256:b8c5aaa9e1a03940f48ba5d970662aa2f2bd540ddde9a02db1cb331c48d0d952

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:0184f8ef302c163089f9d39e4ffab7ae7fefe8389c6fe72bd3045da2857052aa

Observation 304bad2e-afe8-4a0c-86ac-823bd00b71a0 · outbound

This paper cites and Sugiyama, M.

Combinatorial Reinforcement Learning with Preference Feedback and Sugiyama, M

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.224992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.009631Z digest=sha256:302a3901871f209247bbf4ffdbc326e7b63058e90b5b88e37843f1dc1bf188b0

Observation b2ccb11c-209a-402c-ae20-0d9270cbc34b · outbound

This paper cites A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation.

Combinatorial Reinforcement Learning with Preference Feedback A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.024732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.024732Z digest=sha256:9a22249c6350898ae7921112c3080dd2b8bf8024728fcf6a0649c6e5725fe50a

Observation 61a64b14-a5d7-4c2d-91d6-8eb528de124a · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Combinatorial Reinforcement Learning with Preference Feedback Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.158264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.054874Z digest=sha256:ee13126f39d2ddaad32bdb1175a2eb42543a2481e373cd377cf6f6d7e687a0cb

Observation b708c4e8-4c11-4f98-97b7-749ec3e6d313 · outbound

This paper cites Provably efficient reinforcement learning for discounted mdps with feature mapping.

Combinatorial Reinforcement Learning with Preference Feedback Provably efficient reinforcement learning for discounted mdps with feature mapping

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.048206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.080977Z digest=sha256:aebf6af0872fdcc1f3753bebfb86d7ec4efe916928e7004f76560bb0d74e04b1

Observation 1d0be506-c88f-4589-a7c3-2ea55a333c73 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Combinatorial Reinforcement Learning with Preference Feedback Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:13.941327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.108373Z digest=sha256:d169535b5d3124e91e8ec7a3927cf837c9a40c74e66a064eafeebca78b1c14a2

Observation f862ff19-2f46-4f8b-b608-2e8bc716e90a · outbound

This paper cites write newline.

Combinatorial Reinforcement Learning with Preference Feedback write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.158975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.158975Z digest=sha256:bd2b014d4343f7c48007a66f1c4f50ee5c3ad1fbf2a962e3aba49cb652fcbc57

Pith citing papers

Observation 43b8d225-c795-465e-b2a8-95d9a7baee25 · inbound

Improved Online Confidence Bounds for Multinomial Logistic Bandits cites this paper.

Improved Online Confidence Bounds for Multinomial Logistic Bandits Combinatorial Reinforcement Learning with Preference Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:57:04.564210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T19:57:04.381612Z digest=sha256:ca34605e1768a2f857036f9ab90022f24db3d4a12f0207ae16a8326334dea435