Pith. sign in

Paper Citation Record · LEDGER

Combinatorial Reinforcement Learning with Preference Feedback

As of 20 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.10158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10158 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:57:04.381612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T19:57:04.556666Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b18648c4-8d23-482a-b0d4-776f4aea6e69 · outbound

This paper cites Instance-wise minimax-optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Instance-wise minimax-optimal algorithms for logistic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.791208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.224384Z digest=sha256:4d1c34dcd0b72fb074af8898b6b1cba8a012682e93aa7379a234a7ecbaa1aa3a

Observation 28165428-c5d9-4984-b1fe-ef7bca76f6de · outbound

This paper cites Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.761100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.231128Z digest=sha256:143bcc96d864e50b18a17420c5d9949349a53368d44b9afcbc0a320257bd4fc0

Observation 2f556e53-d9b4-478a-93e5-6bad9c169184 · outbound

This paper cites A tractable online learning algorithm for the multinomial logit contextual bandit.

Combinatorial Reinforcement Learning with Preference Feedback A tractable online learning algorithm for the multinomial logit contextual bandit

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.731998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.237577Z digest=sha256:4ff4057d39920de6e976e5f3b78d3a4c26eea1c70ead45b60106fa31fd7de549

Observation 548ffc32-b4db-40e6-afa2-84473a54d7ee · outbound

This paper cites and Goyal, N.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, N

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.706758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.245239Z digest=sha256:cf998f2345f1a24aec55b0ca805300199d52507f42be511d9b19fea22419edd9

Observation c5559935-c99c-4704-94fc-f6e73a4e3652 · outbound

This paper cites Thompson sampling for the mnl-bandit.

Combinatorial Reinforcement Learning with Preference Feedback Thompson sampling for the mnl-bandit

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.681536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.256779Z digest=sha256:64255cd56aa3fb9015522fc873b60df7c7c4d7fe18c6b0f91ba3258870ddd70a

Observation 0080f2ad-5033-432b-92f4-b3169a7009ab · outbound

This paper cites Mnl-bandit: A dynamic learning approach to assortment selection.

Combinatorial Reinforcement Learning with Preference Feedback Mnl-bandit: A dynamic learning approach to assortment selection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.262831Z digest=sha256:2d9b9185e8b686eb339af8394942d4a78857ae4f11e9823fdc5206458bf74efe

Observation 1a6b6990-e2f1-4953-b1bb-3f9504f68356 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Combinatorial Reinforcement Learning with Preference Feedback April: Active preference learning-based reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.271517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.271517Z digest=sha256:5d201cef1d436168fab925f8f604cd6ff21296f555358140cb71eb485ee61165

Observation 6ce715ec-61b0-468f-9f9a-9d88960490eb · outbound

This paper cites and Thrampoulidis, C.

Combinatorial Reinforcement Learning with Preference Feedback and Thrampoulidis, C

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.611297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.284760Z digest=sha256:afd8bc20f9c63ab4afaace45e06e7bc70f357858cab8c757a81f977aa8e6d4b6

Observation 87e0a6b4-94a7-4b9e-af07-8a22e8837bc7 · outbound

This paper cites Distributional off-policy evaluation for slate recommendations.

Combinatorial Reinforcement Learning with Preference Feedback Distributional off-policy evaluation for slate recommendations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.587861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.292863Z digest=sha256:56c5049cbf19a71d874dd92e0be37af2f57af6e9a50f8ceb9661647bdc7a5a1b

Observation 280622c9-3417-4960-854e-bad58b1406bd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.561019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.304989Z digest=sha256:7d736d65f51715dc9eeb9be0bb43701b428f3e6b2978238010792b45f93bcbb2

Observation 2ba47c7e-b279-4d54-8431-699597d2ea37 · outbound

This paper cites Combinatorial multi-armed bandit: General framework and applications.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial multi-armed bandit: General framework and applications

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.532539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.312614Z digest=sha256:4e6eac6de8d1b074b60de8bda3a675c09b60067591783a4cc0ddaa2228be520f

Observation f373dd7e-7ebb-408d-b664-499d8c7959c3 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.507919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.323399Z digest=sha256:a7e3af3bbd5c6f31e4cfae14b1123218803cadad70e3c622c66060ce5074164c

Observation 1301a24b-9616-49e4-be44-00e3b0eebef5 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Combinatorial Reinforcement Learning with Preference Feedback F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.332237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.332237Z digest=sha256:619e8f85c450f588fe2c9979c29d933305ca49d9f48245cb3220a5833952db77

Observation 9f827c34-880e-4eb0-b828-1d6e72e67e0d · outbound

This paper cites S., Proutiere, A., et al.

Combinatorial Reinforcement Learning with Preference Feedback S., Proutiere, A., et al

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.443014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.341524Z digest=sha256:b687daa3d1ccfecf95a870a60e60d0f1ec32a897eb1f2913268b8824bebb7ebf

Observation f486ebc0-5827-49a4-b03e-e77ade3febf1 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.352173Z digest=sha256:4317841a535ad40d1a98f80d90a0dd654e4ed9c1ae9344248d520a588e5af0f3

Observation 6efd8f5d-8ecd-4f79-a205-07b39d9a6b00 · outbound

This paper cites Assortment planning under the multinomial logit model with totally unimodular constraint structures.

Combinatorial Reinforcement Learning with Preference Feedback Assortment planning under the multinomial logit model with totally unimodular constraint structures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.383925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.359505Z digest=sha256:a8d8b4ecd55eb795700c5869004c7bb889cfc2b78558d7686312d54fec05c6f8

Observation d28026fd-a4ba-4d6e-99d7-21851d4eb11f · outbound

This paper cites Reinforcement learning with combinatorial actions: An application to vehicle routing.

Combinatorial Reinforcement Learning with Preference Feedback Reinforcement learning with combinatorial actions: An application to vehicle routing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.337692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.365124Z digest=sha256:e186a526916796ebf5644e464831250c3e301837fccaa9662ae016f6058e8499

Observation dfac6861-d89d-44c1-9e18-c0cfc3117ede · outbound

This paper cites Bilinear classes: A structural framework for provable generalization in rl.

Combinatorial Reinforcement Learning with Preference Feedback Bilinear classes: A structural framework for provable generalization in rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.314252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.371484Z digest=sha256:0e3b6f0715ddc880527b8af34d78740609c38ae7991beb381ec230a733af38d8

Observation ac33eec9-1542-41d9-aa9a-81687450b21a · outbound

This paper cites Cascading Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Cascading Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.850736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.377776Z digest=sha256:52894e322309d7749e575f5dfba7b53ea72a66062c5db389b0a338275ae59a98

Observation cfa91a6f-ffe0-4abc-946c-00885a286f88 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Improved optimistic algorithms for logistic bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.405443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.405443Z digest=sha256:b4490c9d9f056e5a1b95bbd4763f7c47f5048013c062f64bc834c236467997bc

Observation 098c220a-6abd-4050-b907-6c19d2a35cb0 · outbound

This paper cites Jointly efficient and optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Jointly efficient and optimal algorithms for logistic bandits

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.269061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.475614Z digest=sha256:928b682121d4e2671bfcf1bcebdac430a1930cf1428014588ec3eec14623fd3a

Observation 7e340fbd-91da-4e35-ba4b-b239c9866093 · outbound

This paper cites Parametric bandits: The generalized linear case.

Combinatorial Reinforcement Learning with Preference Feedback Parametric bandits: The generalized linear case

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.243063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.546882Z digest=sha256:1874a7fd6c4df367e5c55eab2026440efaa08c21d660a560a1f4796ed8ef6c49

Observation 2f8a65c8-fb8b-4f3f-8849-0f9032f82e28 · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Combinatorial Reinforcement Learning with Preference Feedback The Statistical Complexity of Interactive Decision Making

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.598167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.598167Z digest=sha256:ec0e2e97219b1e8be9fa9fed64a2fea3f0f10c71ffbb608366f6370039fd5bd4

Observation c03b3cd5-d973-41bd-a51c-8dfa139917cd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.218068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.655974Z digest=sha256:89c82be0e966b4954b5c2596ecabc97f4f0b8d3a0efbbd524c4c55d6e9f157fb

Observation 6928cb0d-781e-413a-8e6e-441d96001d7d · outbound

This paper cites Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.748595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.665758Z digest=sha256:5034a7272eeebb3237fcc0a075d69be1e0e06a1a8158fd7bd61b3da60a2b01c9

Observation 3d67be2f-445b-4ee9-8b72-fb62b692493b · outbound

This paper cites Slateq: A tractable decomposition for reinforcement learning with recommendation sets.

Combinatorial Reinforcement Learning with Preference Feedback Slateq: A tractable decomposition for reinforcement learning with recommendation sets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.197624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.672769Z digest=sha256:8e316662670db0b3fba4dbcced74a30675ce19baf4d952dd10439e62d121d176

Observation db054e73-4680-4d27-87cf-20070c8075d5 · outbound

This paper cites Randomized exploration in reinforcement learning with general value function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Randomized exploration in reinforcement learning with general value function approximation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.680375Z digest=sha256:093f3bed604f0dd43694f984969cb9409a1a1bba1ca9c0ccad9fd7d1d47ca3b3

Observation d37ac1b4-fda5-476f-a172-b224f7503961 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.133885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.685870Z digest=sha256:712539feaed276afa328cf43c3843aefa9082a510cce580b008548a4280106db

Observation eb345469-8408-4da7-ab13-a4e73fec3f43 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.102859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.692642Z digest=sha256:99072e9004120be91e0df8fbfb76a539e8036947aaaeeab96c524f5f7f7bc7dd

Observation 768a359f-e859-454d-8b5d-64a5cc65c0f9 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.076408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.699888Z digest=sha256:7c60d7d25a9944ebcaa9c12e3349c355f28a1bc6794228fc757e8b0a8b79ba25

Observation 20505fb5-e7b5-4bb6-bba8-f729ebe59dd9 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Combinatorial Reinforcement Learning with Preference Feedback Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.053824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.705775Z digest=sha256:7745f160ccd1eae275715df14a6a18da6dc8f587842b0565c6b6c13f6603409f

Observation c7559e90-d365-4d5d-887e-5c33a19d1edf · outbound

This paper cites Online Sub-Sampling for Reinforcement Learning with General Function Approximation.

Combinatorial Reinforcement Learning with Preference Feedback Online Sub-Sampling for Reinforcement Learning with General Function Approximation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.711779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.711779Z digest=sha256:83354e319845f3b5565cfb292a558385bcff0634047504fc8f8746e3d72fb697

Observation e51f6d31-7122-487c-b078-3f5864f46902 · outbound

This paper cites Cascading bandits: Learning to rank in the cascade model.

Combinatorial Reinforcement Learning with Preference Feedback Cascading bandits: Learning to rank in the cascade model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.022919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.717594Z digest=sha256:caaa5c100d6bc521d694fec19c6abae75ff788f3900dde7cf791e73ede0c621e

Observation 97c5c92f-6b79-4add-a98d-4cb47ffe6888 · outbound

This paper cites Combinatorial cascading bandits.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial cascading bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.990933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.722642Z digest=sha256:0c7733732977a99087dac2e64db61eb8ad657a4df79aa8f52b27634a20e87918

Observation a616e48a-f068-49d1-b623-93d9fbe07a2e · outbound

This paper cites and Hutter, M.

Combinatorial Reinforcement Learning with Preference Feedback and Hutter, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.969336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.727018Z digest=sha256:f4ed3e7f983e519f5bc0d0ef75d9e9f227d58a69d819c1f097aaa6116ed13202

Observation 1d65c601-5eb0-47a9-a4f8-b27ae658b390 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.946498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.732283Z digest=sha256:712220a5cd52801f1fd53efc9e79877aa4a2e43c97580f69c50ac59dbd0a19c7

Observation 28b636af-c5f1-4f43-896a-d20548e80e13 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.736731Z digest=sha256:6382a5d6e87c7684c6865cf439a67a3b9ef454713bba568d67b59af2dec906b5

Observation aabc5773-222e-482b-97a6-cc8434b97888 · outbound

This paper cites Online learning to rank with features.

Combinatorial Reinforcement Learning with Preference Feedback Online learning to rank with features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.741426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.741426Z digest=sha256:37d08be4bd0b2b83c041e660f94af8b50b4c71cf042a24ed185760e92f2500b8

Observation 2e7da764-f091-4cd8-8af8-a1b6b9e12d03 · outbound

This paper cites Modelling the choice of residential location.

Combinatorial Reinforcement Learning with Preference Feedback Modelling the choice of residential location

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.869998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.747449Z digest=sha256:2b02e7add8ee9900387eaeb45b626f25a001be93a1197343dc7ce1cb30424290

Observation 7ea4aeb3-84a7-439d-8dd2-8f0fb9d3985f · outbound

This paper cites Counterfactual evaluation of slate recommendations with sequential reward interactions.

Combinatorial Reinforcement Learning with Preference Feedback Counterfactual evaluation of slate recommendations with sequential reward interactions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.846788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.754104Z digest=sha256:bd666dc68f114ff5f994b156c31c996f8e0b67f2082469f5aeb6120e214dd60c

Observation fcfff1cd-f4ab-4158-ae4a-b06c3f79319c · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Combinatorial Reinforcement Learning with Preference Feedback Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.761945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.761945Z digest=sha256:058c8c7ff3f4e39015fe070349f813b350645af2d85da6bc2c4ae170bef53758

Observation f28ccee3-682d-4983-af4c-fc167647d0c5 · outbound

This paper cites M., and Van Erven, T.

Combinatorial Reinforcement Learning with Preference Feedback M., and Van Erven, T

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.823133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.792670Z digest=sha256:375f0211839f4592df0bca468861ddd6570205e5cc26b2b2adda8447ccec93eb

Observation 56c2cf88-add8-4503-afae-f42231ec6f72 · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.798459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.825505Z digest=sha256:fc15539c6e8304dbb8d8409229df9342e7145758c372e1a3859d54df1854c73f

Observation 817c122c-11e5-4d5e-a8d0-f173f9c3e92f · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.776371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.844502Z digest=sha256:de4829b1bbb6a49f976d1afffb6ef054480afa0021116c99d5fb6500b0c305d6

Observation 5709ed03-ef9c-4ebf-8579-29b622ec9063 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Combinatorial Reinforcement Learning with Preference Feedback Online Learning: A Modern Introduction Using Convex Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.870065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.870065Z digest=sha256:51cfcf2b07fbbcd3e9722944dd351f8f51213c3755f4b929b7280020d9bd3378

Observation 71557d63-5f85-40df-aeba-389f1c2d1a42 · outbound

This paper cites Training language models to follow instructions with human feedback.

Combinatorial Reinforcement Learning with Preference Feedback Training language models to follow instructions with human feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.896163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.896163Z digest=sha256:f41ac7143a3d7610cb6c85ebd8f93f3774848dfdb37573b7c65d2de5f3f5e875

Observation b44bb0a4-8212-49ba-821f-f13f5e3bc851 · outbound

This paper cites and Goyal, V.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.725451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.904201Z digest=sha256:c1b7c1b499ec35037b0697dc688a5feb6e531fe1796f40dce19b25a999f1b337

Observation a975f62f-39bd-43ad-aa70-cfa62807d193 · outbound

This paper cites M., and Shmoys, D.

Combinatorial Reinforcement Learning with Preference Feedback M., and Shmoys, D

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.702565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.910945Z digest=sha256:e27bdd74d7b7ca42183dc777096a77fc91d30b9fd010ce769cc7110253003c2f

Observation fd73fb0f-6618-47e3-8b24-2fa8b5fc3fcd · outbound

This paper cites and Van Roy, B.

Combinatorial Reinforcement Learning with Preference Feedback and Van Roy, B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.674460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.918859Z digest=sha256:1dda6530d42833580675dca088cc7d4361ec4c2ff209c390da90dff487f2f643

Observation 3b9cb725-dbc2-4f30-886e-b86fde07259b · outbound

This paper cites CAQL: Continuous Action Q-Learning.

Combinatorial Reinforcement Learning with Preference Feedback CAQL: Continuous Action Q-Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.926888Z digest=sha256:8332bf2e22d8a3010432aadd05fdd1b4af5a22c4dd3f433fe5fef7eac477ceb3

Observation c361c8e7-8103-4598-9c9c-54a000858ecd · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Combinatorial Reinforcement Learning with Preference Feedback Dueling rl: Reinforcement learning with trajectory preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.934451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.934451Z digest=sha256:58b32aa031d20562ea020047f6ce5c287a969d92b26e084743d8682e7e7c618b

Observation 0f14d212-1104-4eb3-a749-0100b3ac460c · outbound

This paper cites and Zeevi, A.

Combinatorial Reinforcement Learning with Preference Feedback and Zeevi, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.571717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.943188Z digest=sha256:6a7dfc80aab4902b847450e8c843f6be80e5c11da9f32d540ef7825d14383ceb

Observation e54c55dc-8a4a-4da0-b233-80cb91fa2068 · outbound

This paper cites Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.951927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.951927Z digest=sha256:d409e4ca3ca5942e7245604a349d83328e9959cf28dfd3cdb9a9fba3ac7078ec

Observation 2d6c0153-9fd0-48af-aeb0-753d3caa5d61 · outbound

This paper cites Off-policy evaluation for slate recommendation.

Combinatorial Reinforcement Learning with Preference Feedback Off-policy evaluation for slate recommendation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.490085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.958603Z digest=sha256:edc7a018281ff9fcd2c131b1090346a2e1843032214d19a2b28726f6518cb40f

Observation 966e96de-4d75-4ed4-a76e-7b8f27233eb2 · outbound

This paper cites Composite convex minimization involving self-concordant-like cost functions.

Combinatorial Reinforcement Learning with Preference Feedback Composite convex minimization involving self-concordant-like cost functions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.392696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.967403Z digest=sha256:fa079d896a0acd6a34467406559667cbd25285c2a51c0e360eeffafafee7fd93

Observation 0381e4e0-bdaa-41eb-8fd2-05ade2ab0bc1 · outbound

This paper cites Control variates for slate off-policy evaluation.

Combinatorial Reinforcement Learning with Preference Feedback Control variates for slate off-policy evaluation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.370807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.973733Z digest=sha256:84904ff20326b58811fe1e6bba794c64224df2186b7ca4cf66cb6632a455ace2

Observation 94a32340-6a5b-4085-bb4e-f7ca283eb595 · outbound

This paper cites R., and Yang, L.

Combinatorial Reinforcement Learning with Preference Feedback R., and Yang, L

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.343071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.979355Z digest=sha256:1058cd01b5ae9032c04475a01aca6d1cf593b52a7cddc22e0a721c1a2ea925a0

Observation cf765a08-75ae-4f55-8dba-ffcf813dc963 · outbound

This paper cites S., and Krishnamurthy, A.

Combinatorial Reinforcement Learning with Preference Feedback S., and Krishnamurthy, A

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.299861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.984567Z digest=sha256:407792662f3aff8cd154b0caa42435a4cb02bedab77efd7cc14ca84cd1287879

Observation f7cfa02b-79b3-4be3-9004-935ad45b2f17 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Combinatorial Reinforcement Learning with Preference Feedback A survey of preference-based reinforcement learning methods

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.992315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.992315Z digest=sha256:1d8ddc6b4f7fb937f97dfa449bc09479bd3d12d980e7c08e71aa00eaa3347d5e

Observation 709c594c-281b-4b1f-b016-6bb1a69643b8 · outbound

This paper cites and Wang, M.

Combinatorial Reinforcement Learning with Preference Feedback and Wang, M

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.250244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.998266Z digest=sha256:635dd3c6671dfc66e40cac9f61c2bacd5f0c79ce9d9cc24638d6e3206d7340a7

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:c6b36ec210e1a56dcfe65fc7c00a4a486cf9bfdaa8bc6bc6843a295a7988019e

Observation 304bad2e-afe8-4a0c-86ac-823bd00b71a0 · outbound

This paper cites and Sugiyama, M.

Combinatorial Reinforcement Learning with Preference Feedback and Sugiyama, M

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.224992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.009631Z digest=sha256:48739f2704eb37aa656773a292e750fa69c24b2050bf96d3ae850e8ae178c6fd

Observation b2ccb11c-209a-402c-ae20-0d9270cbc34b · outbound

This paper cites A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation.

Combinatorial Reinforcement Learning with Preference Feedback A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.024732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.024732Z digest=sha256:69e1e62d2af6011ef503836d24e2065dbfd721778087f76d8f34fbe7f16df791

Observation 61a64b14-a5d7-4c2d-91d6-8eb528de124a · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Combinatorial Reinforcement Learning with Preference Feedback Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.158264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.054874Z digest=sha256:5451f89027698d307a5a9316f3114a449d1e4b8fba8f02b238c50bf65a73a55e

Observation b708c4e8-4c11-4f98-97b7-749ec3e6d313 · outbound

This paper cites Provably efficient reinforcement learning for discounted mdps with feature mapping.

Combinatorial Reinforcement Learning with Preference Feedback Provably efficient reinforcement learning for discounted mdps with feature mapping

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.048206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.080977Z digest=sha256:dc22495a9f8fc54010242f9cd8ffc23aeb259ace8bc32a768b69194c6bcd0aca

Observation 1d0be506-c88f-4589-a7c3-2ea55a333c73 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Combinatorial Reinforcement Learning with Preference Feedback Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:13.941327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.108373Z digest=sha256:dfb513686d984451bdb2856ce014615183f4ef3fa174d8daf22d868e76327c2d

Observation f862ff19-2f46-4f8b-b608-2e8bc716e90a · outbound

This paper cites write newline.

Combinatorial Reinforcement Learning with Preference Feedback write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.158975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.158975Z digest=sha256:357212c8f428265b20948e5f385fc79022e7a6fd512193b64785b1c79ea84353

Pith citing papers

Observation 43b8d225-c795-465e-b2a8-95d9a7baee25 · inbound

Improved Online Confidence Bounds for Multinomial Logistic Bandits cites this paper.

Improved Online Confidence Bounds for Multinomial Logistic Bandits Combinatorial Reinforcement Learning with Preference Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:57:04.564210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:57:04.381612Z digest=sha256:f4d676256729ca5463fd34ee15f8c252221f05d589eb041ed11b6f1623a4f05a