Pith. sign in

Paper Citation Record · LEDGER

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.10804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10804 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:27:04.535925Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e219ae47-4f7b-4d24-bdd4-15a4dfab993d · outbound

This paper cites Whittle’s index policy for a multi-class queueing system with convex holding costs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Whittle’s index policy for a multi-class queueing system with convex holding costs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:10.088833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:01.997873Z digest=sha256:3444caf1a31ecaec54f36ef566f268800ed0e3c741b94584d4de2616044c6062

Observation 57191e67-757e-47ab-a81c-20927c5d9c1f · outbound

This paper cites Dynamic allocation indices for restless projects and queueing admission control: a polyhedral approach.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Dynamic allocation indices for restless projects and queueing admission control: a polyhedral approach

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.843742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.074608Z digest=sha256:836225d4675e59f6c1049a88ff577fa3de9c0ce7069f2cbf41a2ebd764e92f1c

Observation 61adbc05-dd6d-496b-83aa-e821ae97042b · outbound

This paper cites Field study in deploying restless multi- armed bandits: Assisting non-profits in improving maternal and child health.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Field study in deploying restless multi- armed bandits: Assisting non-profits in improving maternal and child health

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.581931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.136881Z digest=sha256:de54a5615a421b3deced65291341b425c489e80200a4b493b2d569d00d431dd6

Observation 5634b63f-55e1-4272-afdb-a302923553bc · outbound

This paper cites Collapsing bandits and their application to public health intervention.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Collapsing bandits and their application to public health intervention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.368472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.205380Z digest=sha256:cba5ac1ab444c7b705689fdc9a2be9580640245133a6de67df0fb1107c7e09d8

Observation 36e713a6-2481-4cb5-b13c-0b44269523c4 · outbound

This paper cites Distributed optimal relay selection in wireless cooperative networks with finite-state markov channels.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Distributed optimal relay selection in wireless cooperative networks with finite-state markov channels

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.143708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.269087Z digest=sha256:0bbb3bfd25c40386dbe49fdb9958d80ff50d25a7f4e8d0ebc55c669211df8a26

Observation 6531ab99-6dca-4f7b-bccf-825758c92033 · outbound

This paper cites Cell association with user behavior awareness in heterogeneous cellular networks.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Cell association with user behavior awareness in heterogeneous cellular networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.917036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.338129Z digest=sha256:1feadd112f8e55d259ffb8a2e0f8587f9045b6d4cfed55d1b1e4209214bfc68b

Observation a863df58-7f4a-49c3-a485-c81a1ac39673 · outbound

This paper cites Optimistic whittle index policy: Online learning for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Optimistic whittle index policy: Online learning for restless bandits

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.723592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.408278Z digest=sha256:38a9de2656212c13a32047c17336f6aaf32411a7bbc3ee25f28d299ca874eea1

Observation 489a499a-f7a1-41d8-a302-59a33b08cc3f · outbound

This paper cites Reinforcement learning for non- stationary markov decision processes: The blessing of (more) optimism.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Reinforcement learning for non- stationary markov decision processes: The blessing of (more) optimism

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.424887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.473859Z digest=sha256:6dce4196ac001eaaaefeee2ac0621428639b3197a9ebc8a7b9f55907665e4fc1

Observation 1ef18fa3-6690-455b-8141-7c7e02bd0c39 · outbound

This paper cites Regret bounds for thompson sampling in episodic restless bandit problems.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Regret bounds for thompson sampling in episodic restless bandit problems

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.210494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.566224Z digest=sha256:65ece339e66f4202af929313f31c577e858de0009d1ea00474e5aa59ac17206c

Observation 51d51be1-e07d-41b9-9b0f-d96e31b4ed55 · outbound

This paper cites Thompson Sampling in Non-Episodic Restless Bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Thompson Sampling in Non-Episodic Restless Bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.648338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.648338Z digest=sha256:61ba6eb67962d5770efdbd9b6781c6d1e220324620837fda908500116de0afd9

Observation 95c17267-f2e2-44d0-9843-112aafd5f792 · outbound

This paper cites Restless bandits: Activity allocation in a changing world.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Restless bandits: Activity allocation in a changing world

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.738561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.738561Z digest=sha256:838adafff486fd6c4beb1fc68acb596c2afe136b3615ea48657e7f61a36decad

Observation 1a471b5c-7ae8-408f-aa8e-197778d1fcc4 · outbound

This paper cites Introduction to multi-armed bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Introduction to multi-armed bandits

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.064518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:02.810158Z digest=sha256:430d12909de1808a479c5fbd118574b7d0eccf6c92688c220b47bae9b54622ec

Observation 579c20fe-caee-4318-9d21-6dc70c8324f4 · outbound

This paper cites Algorithms for multi-armed bandit problems.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Algorithms for multi-armed bandit problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.916543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.916543Z digest=sha256:45d23a2b78e123559a9e680b97a7e413cc32594a02bc64ef69711dcd488b4dbb

Observation 1d4e2415-cdb7-434e-a79d-f3ae2fa2cc28 · outbound

This paper cites The complexity of optimal queuing network control.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee The complexity of optimal queuing network control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.998565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.998565Z digest=sha256:3b6e850f6b9ec0f93a8de0ec3bc113533adfd966e5607936aad7a19083d6feb8

Observation 41cbc4cd-3155-420d-bd28-8e757f3872a2 · outbound

This paper cites On an index policy for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee On an index policy for restless bandits

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.814489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.095698Z digest=sha256:3860fa42a7c8d0d3e0f53dd22e558fba19d4e24cb2d9d72b00c35358f38adc78

Observation fa63806b-c50f-40da-9175-ee81a4a4f7e7 · outbound

This paper cites A restless bandit formulation of opportunistic access: Indexablity and index policy.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee A restless bandit formulation of opportunistic access: Indexablity and index policy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.643212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.149894Z digest=sha256:5b827bf52849d15b5237f5355c4cb84818eaac17f384bce9c2eab277cc2aa0b5

Observation 04aad0be-57c6-4ceb-9d86-78f7a112e136 · outbound

This paper cites Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access.IEEE Transactions on Information Theory, 56(11):5547– 5567, 2010.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access.IEEE Transactions on Information Theory, 56(11):5547– 5567, 2010

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.327107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.152865Z digest=sha256:de4934d89173b7396b57fd574df4a89c0633c13011189d05675a613a544a4e22

Observation 3771e6c2-68db-4240-9077-3604d2a8ee23 · outbound

This paper cites Qwi: Q-learning with whittle index.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Qwi: Q-learning with whittle index

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.145281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.186474Z digest=sha256:4d4ae1a7529b54a33cd7923dc5a36c4ca20de0df51525c7b4c26bb403c3763cc

Observation f69f6a6f-b8de-421c-a133-d25ac6296dcc · outbound

This paper cites Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:03.273298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:03.273298Z digest=sha256:8f4d4f276c667be341e74461fe40f48d5d9b85c4471b9680fc34d30f3c333fe5

Observation 96ce33c7-4b11-4ccf-9b88-4ea9e493c82a · outbound

This paper cites Towards q-learning the whittle index for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Towards q-learning the whittle index for restless bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:03.350381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:03.350381Z digest=sha256:b391aa5488dcaf9a8cd2e09fa9b20c409ba8bf26efb32e74f8b5266c49fb19fa

Observation 93a38f69-b969-433e-97bb-cf564ef7329d · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Near-optimal regret bounds for reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.923679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.411313Z digest=sha256:6b85383e97cb4b1f354aaddf664d92b1fcdef72f43224e59de4d3f12a26f117f

Observation 18741a19-acf4-45f8-aa35-a7167068442e · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.669105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.496746Z digest=sha256:6916528bfe0caaa0a9257ed25a39d636c193aac20f55d891fca6dca3a2ccb684

Observation c88f8d83-b79f-41fa-9a46-20deb7bd4c54 · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:05.256092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.662892Z digest=sha256:c42aafffbbe1a890e06466a1caa9f7a425cfb0c505201631456232ecc2a20c22

Observation 5325003f-4bfd-45f6-9c7d-5bfc62280e21 · outbound

This paper cites Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.481285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:03.815378Z digest=sha256:667451639b8e98ddecd48a8f857f80bd531b902fd67f372c324dfde609447751

Observation de96fa00-8325-4043-88b8-9c607e7228b5 · outbound

This paper cites A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:05.038455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:04.035674Z digest=sha256:cbeb31c2328212447c1aeaaff1c8bcac4eb5c9f41672a4d3c3ad04959ebb510a

Observation 0de02645-306b-4cd3-b529-8919694cc5a7 · outbound

This paper cites Near-optimal model-free reinforcement learning in non-stationary episodic mdps.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Near-optimal model-free reinforcement learning in non-stationary episodic mdps

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.110026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:04.203518Z digest=sha256:0762b51f34c17fb52f4a9521281d875081c3ff409568f300a0a891c496e96670

Observation c11553c4-59ca-472e-ab2e-dfb59c929753 · outbound

This paper cites Learning in a changing world: Restless multiarmed bandit with unknown dynamics.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Learning in a changing world: Restless multiarmed bandit with unknown dynamics

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:05.833904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:04.329643Z digest=sha256:8b50f82ded9182dcaea0ce3e2bcd2af0e2aa61e8c2664ddbc0090f995b737f3e

Observation 0ba162fe-673c-4ab2-8097-ecca8b779c80 · outbound

This paper cites Regret Bounds for Discounted MDPs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Regret Bounds for Discounted MDPs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:04.837991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:04.441499Z digest=sha256:5e19b3ab24ec24f35587d2a5b04a134d2dcca7d2fffc0b53f838cee62c3684e2

Observation d81409f8-2fdf-4991-bbee-084a78a9f3b2 · outbound

This paper cites X t∈N γt−1 R(st, at) − λ∗ t πt(st)⊤1 − K | s1 = s # (31) ≤ E(s,a)∼(P ,πt).

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee X t∈N γt−1 R(st, at) − λ∗ t πt(st)⊤1 − K | s1 = s # (31) ≤ E(s,a)∼(P ,πt)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:05.546835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T20:27:04.535925Z digest=sha256:0a4b086f8b21e3a5689ab0aa13fbefc63a0dbb0cf3c75f21042777cb1dbb826b

Pith citing papers

No inbound Pith citation observations are available.