Pith. sign in

Paper Citation Record · LEDGER

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.10804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10804 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:27:04.535925Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e219ae47-4f7b-4d24-bdd4-15a4dfab993d · outbound

This paper cites Whittle’s index policy for a multi-class queueing system with convex holding costs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Whittle’s index policy for a multi-class queueing system with convex holding costs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:10.088833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:01.997873Z digest=sha256:bca524f8bf8e3b545a1e21e3565e02ee68e962b264a2edac3953cbee57bd4fae

Observation 57191e67-757e-47ab-a81c-20927c5d9c1f · outbound

This paper cites Dynamic allocation indices for restless projects and queueing admission control: a polyhedral approach.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Dynamic allocation indices for restless projects and queueing admission control: a polyhedral approach

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.843742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.074608Z digest=sha256:3dbe4264da352915731ac5146611c15fe2cac7313bf708becad893b87c222fe0

Observation 61adbc05-dd6d-496b-83aa-e821ae97042b · outbound

This paper cites Field study in deploying restless multi- armed bandits: Assisting non-profits in improving maternal and child health.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Field study in deploying restless multi- armed bandits: Assisting non-profits in improving maternal and child health

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.581931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.136881Z digest=sha256:5819876836c3b83b7a527807ef9cd71f74d1ab5acffbbbb13d41ec039eb20757

Observation 5634b63f-55e1-4272-afdb-a302923553bc · outbound

This paper cites Collapsing bandits and their application to public health intervention.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Collapsing bandits and their application to public health intervention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.368472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.205380Z digest=sha256:0926c1614f3798dec6fbb382cb738324dbad70dbf3cdc1aa0d3e97c3e095ee2d

Observation 36e713a6-2481-4cb5-b13c-0b44269523c4 · outbound

This paper cites Distributed optimal relay selection in wireless cooperative networks with finite-state markov channels.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Distributed optimal relay selection in wireless cooperative networks with finite-state markov channels

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:09.143708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.269087Z digest=sha256:49e5ed28c90a936fcc062f304972a7f7564985c63a89eddc7d3cebede22474c3

Observation 6531ab99-6dca-4f7b-bccf-825758c92033 · outbound

This paper cites Cell association with user behavior awareness in heterogeneous cellular networks.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Cell association with user behavior awareness in heterogeneous cellular networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.917036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.338129Z digest=sha256:da680fde7fe7541d4b63a5a0e16c761d9cbe8cf60b76422d46687f1dbc771389

Observation a863df58-7f4a-49c3-a485-c81a1ac39673 · outbound

This paper cites Optimistic whittle index policy: Online learning for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Optimistic whittle index policy: Online learning for restless bandits

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.723592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.408278Z digest=sha256:96ae38bb4b0658dfef0102580d12ef6e9673aaf9caeedc8f75ec12d285674f79

Observation 489a499a-f7a1-41d8-a302-59a33b08cc3f · outbound

This paper cites Reinforcement learning for non- stationary markov decision processes: The blessing of (more) optimism.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Reinforcement learning for non- stationary markov decision processes: The blessing of (more) optimism

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.424887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.473859Z digest=sha256:3d2e36da37b77058cdc517b9cca031423182e454508890dff5a7dbb1d2f108de

Observation 1ef18fa3-6690-455b-8141-7c7e02bd0c39 · outbound

This paper cites Regret bounds for thompson sampling in episodic restless bandit problems.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Regret bounds for thompson sampling in episodic restless bandit problems

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.210494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.566224Z digest=sha256:bb7d7a20308928f5136d53241a70d4b124c0ebcbf36f54deecb7bfee5dafff6a

Observation 51d51be1-e07d-41b9-9b0f-d96e31b4ed55 · outbound

This paper cites Thompson Sampling in Non-Episodic Restless Bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Thompson Sampling in Non-Episodic Restless Bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.648338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.648338Z digest=sha256:eec0df03c8ccc4b90bd8bbf6e062059e1bacac0f35f1d9df5385891d53baa3e0

Observation 95c17267-f2e2-44d0-9843-112aafd5f792 · outbound

This paper cites Restless bandits: Activity allocation in a changing world.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Restless bandits: Activity allocation in a changing world

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.738561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.738561Z digest=sha256:547b001f1123c682e088eec11f666342e40bb2153375ed59002258bce8ccdae8

Observation 1a471b5c-7ae8-408f-aa8e-197778d1fcc4 · outbound

This paper cites Introduction to multi-armed bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Introduction to multi-armed bandits

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:08.064518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:02.810158Z digest=sha256:790ac249c1654f68ac1044b459ec47bd2c2652866b094ed3055c9f8efb609447

Observation 579c20fe-caee-4318-9d21-6dc70c8324f4 · outbound

This paper cites Algorithms for multi-armed bandit problems.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Algorithms for multi-armed bandit problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.916543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.916543Z digest=sha256:46192f32dced23d67a0de044d5bb9855bfabef2da8f32becb0e5ce9cddf35c09

Observation 1d4e2415-cdb7-434e-a79d-f3ae2fa2cc28 · outbound

This paper cites The complexity of optimal queuing network control.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee The complexity of optimal queuing network control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:02.998565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:02.998565Z digest=sha256:8878faab8aaee1eb03bc321311797513df85181a09531c4d212cc371bb215ac3

Observation 41cbc4cd-3155-420d-bd28-8e757f3872a2 · outbound

This paper cites On an index policy for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee On an index policy for restless bandits

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.814489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.095698Z digest=sha256:3b8513f56f45f0671a5948f1cbf286f00b8fb616b47e5c191135a522a47032fb

Observation fa63806b-c50f-40da-9175-ee81a4a4f7e7 · outbound

This paper cites A restless bandit formulation of opportunistic access: Indexablity and index policy.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee A restless bandit formulation of opportunistic access: Indexablity and index policy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.643212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.149894Z digest=sha256:8582a57fa0dbe5b2a08b9d565ffdb4001a5214410e2a802b78f265b392138b0a

Observation 04aad0be-57c6-4ceb-9d86-78f7a112e136 · outbound

This paper cites Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access.IEEE Transactions on Information Theory, 56(11):5547– 5567, 2010.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access.IEEE Transactions on Information Theory, 56(11):5547– 5567, 2010

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.327107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.152865Z digest=sha256:83335590ed64d4df93230a5a7e4394398fece1c746afbdd18c0dc189955ad389

Observation 3771e6c2-68db-4240-9077-3604d2a8ee23 · outbound

This paper cites Qwi: Q-learning with whittle index.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Qwi: Q-learning with whittle index

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:07.145281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.186474Z digest=sha256:5f458c9307f1795d71ea833405c3d118b3c3f692f42b31a77dfcb40c8aa78cb2

Observation f69f6a6f-b8de-421c-a133-d25ac6296dcc · outbound

This paper cites Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:03.273298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:03.273298Z digest=sha256:c0482262b257fe2dfa8bb3d3825517109afe62e7668e72dbcb19822c7a93c6a6

Observation 96ce33c7-4b11-4ccf-9b88-4ea9e493c82a · outbound

This paper cites Towards q-learning the whittle index for restless bandits.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Towards q-learning the whittle index for restless bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:27:03.350381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:27:03.350381Z digest=sha256:e635305ada33a522601adc066aa92c03ee800d218d2116c6f8c910f396ac258a

Observation 93a38f69-b969-433e-97bb-cf564ef7329d · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Near-optimal regret bounds for reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.923679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.411313Z digest=sha256:b7744c4d6241b6bf434c18ec9c47bb7a3ebdf63a7801289fc93fa35790b27361

Observation 18741a19-acf4-45f8-aa35-a7167068442e · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.669105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.496746Z digest=sha256:3aab8404f0e3ce5dd31808507295dc9f3beebb84ab0f409a5d687d2fa79d8130

Observation c88f8d83-b79f-41fa-9a46-20deb7bd4c54 · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:05.256092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.662892Z digest=sha256:19d4d32f56e4d08bfcafe493617a487bb54109c0f51a4383c6b02a85e9b4dc7a

Observation 5325003f-4bfd-45f6-9c7d-5bfc62280e21 · outbound

This paper cites Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.481285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:03.815378Z digest=sha256:0d8138023a271372e40db5d55cadb0baa684aa82a1e6deed518ed73563729f81

Observation de96fa00-8325-4043-88b8-9c607e7228b5 · outbound

This paper cites A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:05.038455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:04.035674Z digest=sha256:3a75df795b8304a2c7acc92a160cf9772da100da62e240c6f4334bccfa4f371c

Observation 0de02645-306b-4cd3-b529-8919694cc5a7 · outbound

This paper cites Near-optimal model-free reinforcement learning in non-stationary episodic mdps.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Near-optimal model-free reinforcement learning in non-stationary episodic mdps

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:06.110026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:04.203518Z digest=sha256:5a34db256024c9783677681d7aabcb817067cdb6b8bbe89e8a881b7c3072fee9

Observation c11553c4-59ca-472e-ab2e-dfb59c929753 · outbound

This paper cites Learning in a changing world: Restless multiarmed bandit with unknown dynamics.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Learning in a changing world: Restless multiarmed bandit with unknown dynamics

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:05.833904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:04.329643Z digest=sha256:d9acdd6db8eeb73d3a253bb50ad6699cbdf8cc22652d05cafcee512a63c2b8be

Observation 0ba162fe-673c-4ab2-8097-ecca8b779c80 · outbound

This paper cites Regret Bounds for Discounted MDPs.

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee Regret Bounds for Discounted MDPs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:27:04.837991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:04.441499Z digest=sha256:604062902762193a7505a00f8b417947699ddf532a4291b85ecb047d294fa5c1

Observation d81409f8-2fdf-4991-bbee-084a78a9f3b2 · outbound

This paper cites X t∈N γt−1 R(st, at) − λ∗ t πt(st)⊤1 − K | s1 = s # (31) ≤ E(s,a)∼(P ,πt).

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee X t∈N γt−1 R(st, at) − λ∗ t πt(st)⊤1 − K | s1 = s # (31) ≤ E(s,a)∼(P ,πt)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:27:05.546835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:27:04.535925Z digest=sha256:226434f890af1188f5204ec57c00daac0ab7b8a6c95e0d0450024f33a3350e87

Pith citing papers

No inbound Pith citation observations are available.