Pith. sign in

Paper Citation Record · LEDGER

Contextual bandits with entropy-based human feedback

As of 10 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2502.08759.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08759 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:51:35.789930Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T05:06:05.369793Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.303918Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b329e6c-839b-45aa-be4a-452e93697d19 · outbound

This paper cites GPT-4 Technical Report.

Contextual bandits with entropy-based human feedback GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.699578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.699578Z digest=sha256:6b1bfb2ebc6eb656af3f86a055f3fce981dbf32d83b8131041ca9b5a67acd1b3

Observation f42ea444-bab5-46ff-b8b5-c3776a0f9f52 · outbound

This paper cites Survey on appli- cations of multi-armed and contextual bandits.

Contextual bandits with entropy-based human feedback Survey on appli- cations of multi-armed and contextual bandits

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.257362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.728588Z digest=sha256:66b6e152fd74981380f5870977cd9b7ee8f94cc535fc2ccff0e9de5f75027185

Observation 043df074-a09e-41f9-ade2-37d263d88070 · outbound

This paper cites Thompson sampling with the online bootstrap.

Contextual bandits with entropy-based human feedback Thompson sampling with the online bootstrap

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.746581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.746581Z digest=sha256:79fd8b77f4c42801924c6c8dab16e9dbf3dcc68036d19c0556f8d85e79d14012

Observation 60c0043b-797e-44af-929d-60620289ba07 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Contextual bandits with entropy-based human feedback Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.756299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.756299Z digest=sha256:6f808ea770205b65147698516adf2a47d73f601095f2b830dd0015794144ac9e

Observation 6f654b1b-2411-418f-976b-9e2e4d522bcc · outbound

This paper cites and Lefebvre, S.

Contextual bandits with entropy-based human feedback and Lefebvre, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.227895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.765013Z digest=sha256:4edc1c498dc3679bc02ea69e439afbb5263020151de6667c3eb2ad9d07304e99

Observation 9a2641ea-dd53-4d65-83f8-4d8838d5908f · outbound

This paper cites Borda Regret Minimization for Generalized Linear Dueling Bandits.

Contextual bandits with entropy-based human feedback Borda Regret Minimization for Generalized Linear Dueling Bandits

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:51:35.866329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.769075Z digest=sha256:a1bfddeb8b62f5d230b9efa1c2f0e4b4d1ec636df772cfb69ba94a5f8934fa39

Observation 3eb569e2-74e1-4e2e-b5be-a6537b2ea515 · outbound

This paper cites FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback.

Contextual bandits with entropy-based human feedback FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.773348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.773348Z digest=sha256:a6535d74fceffffd3083906c22be4cd84d943f5fe77ad9056c03274a71ae7964

Observation 68945443-7639-49d9-a6f4-53eb5da189de · outbound

This paper cites CAREForMe: Contextual Multi-Armed Bandit Recommendation Framework for Mental Health.

Contextual bandits with entropy-based human feedback CAREForMe: Contextual Multi-Armed Bandit Recommendation Framework for Mental Health

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:51:35.829846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.778358Z digest=sha256:6e399909fab960655f5491c6afe08ff9b35312205f9b1199d729a51547c61929

Observation b027f9b1-4418-4521-88de-50160e7142ea · outbound

This paper cites These connections highlight how our approach advances real-time feedback integration and decision optimization.

Contextual bandits with entropy-based human feedback These connections highlight how our approach advances real-time feedback integration and decision optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.185489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.789930Z digest=sha256:cc0ad4d2193515f800baa7fffe0a43d431b2638ec59f53f651a2994ff2eec76e

Observation 0f97e972-8758-4f94-8edd-8bf40a2ae367 · outbound

This paper cites EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits.

Contextual bandits with entropy-based human feedback EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.714690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.714690Z digest=sha256:09eded5ab95f30292526d41d759ed15e70929da8b459a76ec64991ba18a9677a

Observation e5f631b0-48e9-4dd1-8bde-3ecdd832eefd · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Contextual bandits with entropy-based human feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.751446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.751446Z digest=sha256:b970d76095bdc7323e6e466173b0b0be8fa4a7d79d5701b998eb6494c38b5d91

Observation b1fc2aaa-78a2-40a4-ad5e-db4c9de09e94 · outbound

This paper cites an unresolved cited work.

Contextual bandits with entropy-based human feedback Unresolved cited work

Reference 2011

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:51:36.199877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.786303Z digest=sha256:4a71616891c160c98c5642564b0c4c9aafc20ee34914d32b34f96349b43599e6

Observation 924ee04e-f949-43f0-befb-3c9a2f7be854 · outbound

This paper cites A neural networks committee for the contextual bandit problem.

Contextual bandits with entropy-based human feedback A neural networks committee for the contextual bandit problem

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.286247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.704685Z digest=sha256:21e66c2e3a924080e4a63b7169956df0bf1c011098571cf87b2ce959b64c549d

Observation 7c2b461e-b3f5-4629-ad64-d5ed89b25720 · outbound

This paper cites DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback.

Contextual bandits with entropy-based human feedback DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.709372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.709372Z digest=sha256:0d982764c74e6919ec3ccdaadf03173f362c6a8bbc0ed2c77dfab2a14d5ad1f2

Observation dcd5ba6d-5710-4445-935e-463beb14eb02 · outbound

This paper cites Bayesian Active Learning for Classification and Preference Learning.

Contextual bandits with entropy-based human feedback Bayesian Active Learning for Classification and Preference Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.741570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.741570Z digest=sha256:dcb29db3744935fde773bf4f1127745a333472df5e849a1fc250ec2871166beb

Observation da1357dc-0c76-444e-a6f2-231d2566678c · outbound

This paper cites Contextual Bandits and Imitation Learning via Preference-Based Active Queries.

Contextual bandits with entropy-based human feedback Contextual Bandits and Imitation Learning via Preference-Based Active Queries

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.760624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.760624Z digest=sha256:2234b8a056ca6dd6b3f35e5a53ba4aaa59d4a93ebff695681756f594a65439ad

Observation 1463b229-62a1-4005-a5c0-993563767aef · outbound

This paper cites an unresolved cited work.

Contextual bandits with entropy-based human feedback Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:51:36.241736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.732765Z digest=sha256:771b507a5c3c6f33b468d79df5c634f836b35bcf945b93bcbb909212e6e8058d

Observation 8f0b24d6-f70e-4d5f-bc2c-b702d3fc6260 · outbound

This paper cites We draw inspiration from Tang and Wiens (Tang & Wiens, 2023), whose counterfactual-augmented importance sampling informs our feedback framework, and extend DAGGER (Ross et al.,.

Contextual bandits with entropy-based human feedback We draw inspiration from Tang and Wiens (Tang & Wiens, 2023), whose counterfactual-augmented importance sampling informs our feedback framework, and extend DAGGER (Ross et al.,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.214241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.782304Z digest=sha256:4077f5af4919ced33b1c5db461db66f5324971c4a340714f1ace7b29ead58e71

Observation 8defbab9-48ec-4a59-985a-af04f1c8837e · outbound

This paper cites Adversarial Rewards in Universal Learning for Contextual Bandits.

Contextual bandits with entropy-based human feedback Adversarial Rewards in Universal Learning for Contextual Bandits

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:51:36.125117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.719330Z digest=sha256:e8d774f0089db94f29c81c9bcd1ee37a6f2ebcba0ea4e3b5afe7fa6fd5438d36

Observation af5bd6da-e7f1-4604-aa21-281231ee939f · outbound

This paper cites Contextual bandit for active learning: Active thompson sampling.

Contextual bandits with entropy-based human feedback Contextual bandit for active learning: Active thompson sampling

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.272001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:51:35.724103Z digest=sha256:1bbfd565dd8ea0e6124cd72b243ddb6fe34d0832d881dfed6d782c621e310e95

Observation 41d1a8fb-302e-4ddf-8cde-1e20dcaff48e · outbound

This paper cites Nearly optimal algorithms for con- textual dueling bandits from adversarial feedback.

Contextual bandits with entropy-based human feedback Nearly optimal algorithms for con- textual dueling bandits from adversarial feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.737043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.737043Z digest=sha256:11ecfdef1186f12c1a73cc5dde1ce8edcda728cce50ca77d34a9f0e5e9fc860e

Pith citing papers

Observation 15283c2d-6fbd-4d8f-bf12-b5bff76920e6 · inbound

Autoformalization of Agent Instructions into Policy-as-Code cites this paper.

Autoformalization of Agent Instructions into Policy-as-Code Contextual bandits with entropy-based human feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.305589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:06:05.369793Z digest=sha256:f8309f6653a1becf985345eaeb82ed46f1ce9e734aa64d8fce47ff209e3c7395