Pith. sign in

Paper Citation Record · LEDGER

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework

As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.15856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15856 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:34:32.435835Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f531a07-aa77-45b2-b60b-98872ecd42d3 · outbound

This paper cites Threshold bandits, with and without cen- sored feedback.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Threshold bandits, with and without cen- sored feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.962137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.291625Z digest=sha256:92d97787c89aa894cbf86b5effa7ecf018fe3842e3bd6fc50cf928ca77b11a8f

Observation e3315ede-c7fb-4eaa-813a-db83957c31df · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Asymptotically efficient adaptive allocation rules

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.759844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.370030Z digest=sha256:fd05ecc26ae72b0fbe85bece92d091d611da3b2cb82724d50f6289f477c1391c

Observation 8665debe-a4a6-4909-9f16-b3ce5c53a68b · outbound

This paper cites On distributed cooperative decision-making in multiarmed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework On distributed cooperative decision-making in multiarmed bandits

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.717408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.381352Z digest=sha256:d86a98784537eb372f51d93b8222fefc2c3a5d8468aeb0783ae8ae2d97702860

Observation b747bd04-cc0f-4a93-b293-898c383c476c · outbound

This paper cites Social imitation in coopera- tive multiarmed bandits: Partition-based algorithms with strictly local information.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Social imitation in coopera- tive multiarmed bandits: Partition-based algorithms with strictly local information

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.699203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.386557Z digest=sha256:7e7ba7dea249aec820f5543f21919df2527b2dde209b65fc4ec9d46369acf3a7

Observation cd592c1f-edc7-412c-a330-3b2c5d319825 · outbound

This paper cites Bandit algorithms.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bandit algorithms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.662968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.397080Z digest=sha256:d62640627623d4c2d87b44ea5cfbec6115453ada48719c6f1499105a29af5a8e

Observation 5c40fdaa-0b59-461e-84b6-268bd0b0a603 · outbound

This paper cites Decentralized Cooperative Stochastic Bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized Cooperative Stochastic Bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:34:32.413143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:34:32.413143Z digest=sha256:554ff10315ad7c4012ead922408a7ef9efd3de89b3775deba96ced0b06cda3ab

Observation 49cdaa70-c173-42a8-b78c-6b395fcafe73 · outbound

This paper cites Multi-player bandits–a musical chairs approach.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-player bandits–a musical chairs approach

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.602166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.418683Z digest=sha256:4d0b34c4e51e02005ba92f53fb6bbfd93e85e041b11298a6d94117a4152a7844

Observation d14454d1-750a-49ed-9977-7134ddc88cbd · outbound

This paper cites Balanced and incentivized learning with limited shared information in multi-agent multi-armed bandit.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Balanced and incentivized learning with limited shared information in multi-agent multi-armed bandit

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.583749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.424466Z digest=sha256:d70d15a828b578b65a3d549591bca1a20d43e37b222f46a5523d12ed9601961e

Observation f02470dc-7acb-4bc4-843f-e619ae8530c0 · outbound

This paper cites Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:34:32.496701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.435835Z digest=sha256:57ef158d8474804baf1e29ec129b7fb987a1d324d36b5b1f95885bde7b56976f

Observation 8ea434fa-7a54-46c3-bbbb-cf86ae413172 · outbound

This paper cites Decentralized learning for multiplayer multiarmed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized learning for multiplayer multiarmed bandits

Reference 1979

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.777775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.363727Z digest=sha256:71a3ad3b51f7faecbd504eff7c564d6bf71d6e6fa8bf81b7d4a620183c2fe0e2

Observation fb96419b-bb2d-4651-8f75-64a72f6de0ff · outbound

This paper cites Bayesian algorithms for decentralized stochas- tic bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bayesian algorithms for decentralized stochas- tic bandits

Reference 1985

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.737730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.375617Z digest=sha256:ff6b315e38766a663da511ecb6f7a0210abb0d23bda7e9382f983edf71e2ba78

Observation ba844463-ca94-48d4-9eff-2e95db99653a · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Finite-time analysis of the multiarmed bandit problem

Reference 1995

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.908689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.310131Z digest=sha256:156f4653682f90259bc60774d92b32f0b33194f48ed0759822dc91666524774b

Observation c09d8a75-f1d4-42f5-9713-31f83bdecde9 · outbound

This paper cites Communicating with unknown teammates.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Communicating with unknown teammates

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.889492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.315741Z digest=sha256:8fa247ea1749639044dfce662388b4e7ab91049eaa2caf61444255d79984210e

Observation 6a4e4c08-f5c3-4fac-83ca-ef26f51c5649 · outbound

This paper cites Decentralized heterogeneous multi- player multi-armed bandits with non-zero rewards on collisions.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized heterogeneous multi- player multi-armed bandits with non-zero rewards on collisions

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.621530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.408346Z digest=sha256:1f9c7123b2f437d83d8162a24c53308ba8182f50158e32cbfcd3ea7c8ac2388f

Observation 8f2f8cbd-1210-40cf-af60-0ef5947bc0e0 · outbound

This paper cites Coordinated versus decentralized exploration in multi- agent multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Coordinated versus decentralized exploration in multi- agent multi-armed bandits

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.816167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.345834Z digest=sha256:148bac397019edcabd2c0d6427e49dbfed758e4818d0da38e43be5bf2e0c45bd

Observation b08c4370-9e0b-49cf-8dca-188470a98880 · outbound

This paper cites Game of thrones: Fully distributed learning for multi- player bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Game of thrones: Fully distributed learning for multi- player bandits

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.872045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.321499Z digest=sha256:61c6cb7856bb56d567c07e00b2864ddad77207870ee9f3753b27cbedd9bc7b43

Observation 5b6723f9-ee7f-4168-a16f-7adedc730406 · outbound

This paper cites Multi-agent multi-armed bandits with limited communication.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-agent multi-armed bandits with limited communication

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.945226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.298686Z digest=sha256:be7158b2eb9bf6cfa95500920802bf2ed50e3fd97a85b5cb12d9178bda9a1e32

Observation bf4a7aca-b7d9-4bdf-874d-60d88e9f27d5 · outbound

This paper cites Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:34:32.352464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:34:32.352464Z digest=sha256:835d7671fcf32614e7bda18808fb8dbedd2f9fc559c38c31c9b37a3b16b108ca

Observation 47774e1c-d879-4c1c-aff5-dd156cc021d3 · outbound

This paper cites Distributed cooperative deci- sion making in multi-agent multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Distributed cooperative deci- sion making in multi-agent multi-armed bandits

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.680991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.392058Z digest=sha256:aa1ba098ea30571bff01eed670a584368cea6b76ea45fddecad2fc60393feb73

Observation 90557fa2-6e8d-48d5-8601-8cda7787345a · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.836025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.335859Z digest=sha256:3460a6f6288e9f873e392ea04417b0c01e9764fe0c6fde3be99ce7264a93c94f

Observation d3d053b0-c22d-45ca-89d6-a455e00d2949 · outbound

This paper cites Distributed learning in multi-armed bandit with multiple players.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Distributed learning in multi-armed bandit with multiple players

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.642473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.403849Z digest=sha256:5362e8c396b577d92a1e0adefbfe484f129405c28045968d0abe96459587558d

Observation 91536ef5-ece8-4479-ab15-46bfc395934a · outbound

This paper cites Sic-mmab: Synchronisation involves communi- cation in multiplayer multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Sic-mmab: Synchronisation involves communi- cation in multiplayer multi-armed bandits

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.854327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.329919Z digest=sha256:550390f384a00a7ca57daba756add63ea6e2eca23b581038a997fbba9d39ccf3

Observation a34af3d0-ca81-4a36-9627-b9c04d7eb125 · outbound

This paper cites Gambling in a rigged casino: The adversarial multi-armed bandit problem.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Gambling in a rigged casino: The adversarial multi-armed bandit problem

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.926264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.304700Z digest=sha256:0927c202f422849d11f91fe6a21b9a858c942021288f6f74dc10f8699375a31c

Observation cc8474a5-2653-49c6-9c2c-a885495b03b1 · outbound

This paper cites Bandit processes and dy- namic allocation indices.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bandit processes and dy- namic allocation indices

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.798230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.358412Z digest=sha256:9537fc7f5410b20fe935d54e887e4b13e5e8b9cf671eb0743829e64ca659ee81

Observation c3dfcf2b-268f-495b-9187-afdcc8ccd723 · outbound

This paper cites Ad hoc autonomous agent teams: Collaboration without pre-coordination.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Ad hoc autonomous agent teams: Collaboration without pre-coordination

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.564075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:34:32.430106Z digest=sha256:fb079ae07489fcf20442fc1b33aa50fcb9a9f4deca0ba2132d7d478cdd9a3f2c

Pith citing papers

No inbound Pith citation observations are available.