Pith. sign in

Paper Citation Record · LEDGER

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2501.13394.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13394 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:18:17.535895Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact5
  • verified fuzzy34
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02eb01fd-2b88-409d-b363-75ffe4affd28 · outbound

This paper cites write newline.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.332912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.332912Z digest=sha256:1cae4c37d9596f8b782b650d6e77033db6fffc20eeaf18420fa7b01fcf2b7596

Observation 0099de47-d04d-4054-af78-4e0ff8b8a05e · outbound

This paper cites H., Chalup, S.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration H., Chalup, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.199357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.337507Z digest=sha256:246421887d340f18824425443c907bba7194842edebd0a1245a8b28f829ae78b

Observation bc7267e0-51fc-4c9c-8901-70f6b521745d · outbound

This paper cites Making Contextual Decisions with Low Technical Debt.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Making Contextual Decisions with Low Technical Debt

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:18:17.762973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.341200Z digest=sha256:ab500d2cbd4e43fda358089fd793c0d2e8f8cd8607b8e27b59019e73c0ed8ce4

Observation dcc3fb88-bbae-485f-8aa7-92464c32d0e5 · outbound

This paper cites M., Lee, J.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration M., Lee, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.189991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.345338Z digest=sha256:d88c618ec9d033427da1444cbfe9fd0d429bb3ffd17359cb765dfb8a3d550eb9

Observation 5782fdc6-b790-4eef-8a3d-6007ed2115e4 · outbound

This paper cites Improved worst-case regret bounds for randomized least-squares value iteration.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Improved worst-case regret bounds for randomized least-squares value iteration

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.180779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.348745Z digest=sha256:80c033b9a2e054a8aa2a588c1dcaff64770f48e152ea2d0b4c4c2d185209db1f

Observation 4f61e943-ee00-489e-9aaa-7b68e1a8f149 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Near-optimal regret bounds for reinforcement learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.352216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.352216Z digest=sha256:19d9f3aa11bb85a09f2b26816b77246d9bbf97bc04de0e8b7d58b381fd065e48

Observation 68359b24-00d1-4ac5-92fc-e636ee94386f · outbound

This paper cites G., Munos, R., Ghavamzadeh, M., and Kappen, H.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration G., Munos, R., Ghavamzadeh, M., and Kappen, H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.165845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.355443Z digest=sha256:79e67da9e8e93af182d3becbca31a3f4f9c4ec3a2b90169510cbb70eeb1f207e

Observation d29dac7b-09fe-40c3-ad64-6851b57972ed · outbound

This paper cites Emergent Complexity via Multi-Agent Competition.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Emergent Complexity via Multi-Agent Competition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.358850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.358850Z digest=sha256:d8dd1d6440345c502e81df9fdddd30fea096dbe26e894cd7b9b6a1242230c8f5

Observation c9c514a6-8f8d-4fec-a650-1e9eea0d3f66 · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.362465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.362465Z digest=sha256:879d32dc19aeef0d9efc8ca2684b16183ad84102993a9375d636a824cb7c4790

Observation 823456aa-af2e-4d52-804f-95c46703b470 · outbound

This paper cites Multi-agent reinforcement learning: An overview.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Multi-agent reinforcement learning: An overview

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.156934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.366014Z digest=sha256:93ef2b712e8f653380151ef248bf201d9d84d067a47b0f34610a8c09d9fc53dd

Observation 38817715-be49-4ea1-b7c0-cc3ec889da31 · outbound

This paper cites Society of agents: Regret bounds of concurrent thompson sampling.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Society of agents: Regret bounds of concurrent thompson sampling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.147754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.369418Z digest=sha256:c1af2b72b23b8534c01da882e9dbe04074f1bd5c7418164d6c83f328f7adbf8a

Observation 452dc6ca-3001-4a22-bcd9-9255736d7a60 · outbound

This paper cites A provably efficient model-free posterior sampling method for episodic reinforcement learning.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration A provably efficient model-free posterior sampling method for episodic reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.138606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.372932Z digest=sha256:f27a927e8569d8d33b6457d16c2c41189ce1501a54c213dd7f2c9a1e6ce36416

Observation 9d9ec6c9-c161-46b4-96ff-5db588d2ef22 · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:18.129176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.376237Z digest=sha256:4d1f8cdce6f173669094e9e691d51d6228b2c1854bb0818eb0207a6f1d5843aa

Observation 78d5a7a3-a4e5-434a-974a-07e4a671e27c · outbound

This paper cites and Van Roy, B.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Van Roy, B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.119587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.379428Z digest=sha256:a83b7b09d3dc495af0a574b121120cbcb70d3ab50b3bed86b46e01dc7edd1a4c

Observation dbb5718e-7d74-47db-bd6d-42b098e06ad7 · outbound

This paper cites Scalable coordinated exploration in concurrent reinforcement learning.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Scalable coordinated exploration in concurrent reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.110347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.382882Z digest=sha256:3bfa70a6b34b6046a9de74a8c0f0f34ec70e92ab141a4c9f8221fd8a3956ec85

Observation e73cbf13-dca9-4c8c-b6d9-743b092cd34e · outbound

This paper cites Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.386389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.386389Z digest=sha256:41ff4bfc41396697953c494f8bdcb6ac9be452cf8000903f42b3c322f58ae863

Observation 6b58b1e7-c05e-43d9-bd33-fbb2f0f2c266 · outbound

This paper cites Provably Efficient Reinforcement Learning with Aggregated States.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Provably Efficient Reinforcement Learning with Aggregated States

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.390316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.390316Z digest=sha256:ada55d1c197051bffb30b3fe840af8fc00b4e92b50200290964ba8bc8ad3a459

Observation 51f45872-1e53-4662-9974-3a31d97f8e04 · outbound

This paper cites Simple agent, complex environment: Efficient reinforcement learning with agent states.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Simple agent, complex environment: Efficient reinforcement learning with agent states

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.100341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.394064Z digest=sha256:52f989f40b8a880b81b81760514ae6ab7e9a905a2c7bbb01c2be68a7347da961

Observation 8d3f16d0-dfa2-4eeb-9ab3-ca1f455cff8c · outbound

This paper cites Provably Efficient Cooperative Multi-Agent Reinforcement Learning with Function Approximation.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Provably Efficient Cooperative Multi-Agent Reinforcement Learning with Function Approximation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.397466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.397466Z digest=sha256:307f1d02ee64cb8383d2d74e35bb282dae9ecf19ab8622f252dd299ae68797c8

Observation b4f6bb5b-e54b-46a1-b1fd-8f708fa31cae · outbound

This paper cites Hypermodels for Exploration.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Hypermodels for Exploration

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:18:17.702689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.401271Z digest=sha256:4447edcb2310820f4f9ca106f0a92885f5dbdff267f5db85a09024aab9ee4778

Observation 8cada472-bec8-42f6-809f-683292db9677 · outbound

This paper cites Bayesian bellman operators.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Bayesian bellman operators

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.089595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.404910Z digest=sha256:1ab7129884bd049ee29c65d55a6b995cfa8e18013d8de0b139ceee943e048fd2

Observation 3963d48d-2ad9-4107-bcab-c9eb78fb4c3a · outbound

This paper cites Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.080301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.408275Z digest=sha256:4a1e46d860db2f52c173f3434464403b01ffe7919a6c9215e21e129419a08fdf

Observation 31e57f91-8593-4f52-87bf-0f91b49035b7 · outbound

This paper cites and Brunskill, E.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Brunskill, E

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.070456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.411736Z digest=sha256:2a7b430b68720c31c9c78f9a958016f2bcb8f2bbe4e6a24ef09b515bb3cd5189

Observation 06d4e5e6-a1ae-41c1-86e2-3f48e1a06125 · outbound

This paper cites Randomized exploration in reinforcement learning with general value function approximation.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Randomized exploration in reinforcement learning with general value function approximation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.060969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.415159Z digest=sha256:c496d4af1d462b5dba902829f0931b544cd02526f236638dc3f8bf2dc4144a8e

Observation 8d903ea8-6c0a-4264-a522-401f8c096966 · outbound

This paper cites Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.418546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.418546Z digest=sha256:4e5814c8776c16871be11d73e34f20d58ef64146dd22caeb0027b41ad47798ac

Observation b87360a2-93b8-400e-98b9-56edf27c755d · outbound

This paper cites M., and Tschiatschek, S.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration M., and Tschiatschek, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:18.051084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.421996Z digest=sha256:1acb1832ce721a4d4f9a031ccf46ed5fb6cd4f9fbd12fb005e091cbb54aeca5a

Observation de87b722-3feb-448a-9597-7df7c424a35f · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.425392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.425392Z digest=sha256:f9c1841adf281f1e8565616e890be84fe9c2b13f3b33d07694edddb770ac6926

Observation 19ff4a63-3073-49a5-aecd-d1b9d64f2996 · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.429589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.429589Z digest=sha256:4d1e370112df49b9c551f31965620924575466107f659f5215ca7ea5351966a7

Observation 74563cfa-979c-4edb-91bc-ebf2227c29bc · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:18.029056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.433136Z digest=sha256:7d03f6eff3f8e988ebef6ec765a99191874c22800d6daa68cdb585ecb498b52f

Observation 13d21b87-8b4d-4b15-af11-bac77edc4924 · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:18.018562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.436300Z digest=sha256:aac9a61edc5edc9e4cdc8d864b62f41c77a015e1696756c9073a8b38c9e79c39

Observation 2e967ad5-0af7-4832-ba2a-cfcf81017074 · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:18.008512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.439544Z digest=sha256:954eb7891a6666f2ca0a0b6e8f16ce11e9574097610475e4a791693ac6362363

Observation a3288c5d-451a-4f9c-95da-3a726153e7ce · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:17.997834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.442674Z digest=sha256:738f70218e38faa049befc27bbc0abc39a2e35dc1a142fdffca50f97ab56b33f

Observation 1db6888e-1b75-44a5-8019-4fe0a6ac4515 · outbound

This paper cites I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.445834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.445834Z digest=sha256:a0bcd4bc9528dbd65b6acfd86164f9e3eecea6949320460163374572b7722ce4

Observation 018d260c-4c7e-4756-b50c-4a00861a092d · outbound

This paper cites Cooperative multi-agent reinforcement learning: Asynchronous communication and linear function approximation.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Cooperative multi-agent reinforcement learning: Asynchronous communication and linear function approximation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.448976Z digest=sha256:db6bcbf4f4c12f761da95ec0c5d48140f24c9cbd438cf02cad174b63fecb6348

Observation d56e2646-40ab-45f0-a543-b97324af78de · outbound

This paper cites and Van Roy, B.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Van Roy, B

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.971547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.452154Z digest=sha256:13802e8861e6053c27ba03af2b3a10ab843d31451aa6bbec4b38839dbfc96405

Observation 7af09584-591b-402a-9623-62dea23801c6 · outbound

This paper cites On Optimistic versus Randomized Exploration in Reinforcement Learning.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration On Optimistic versus Randomized Exploration in Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:18:17.675088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.455424Z digest=sha256:e8795d8f70b00a2c12d98999a8a9cbb69a98dadb5238a1eadeefd7b94c55762b

Observation d8e80065-e474-4124-8c97-84c31898d9e2 · outbound

This paper cites ( M ore) efficient reinforcement learning via posterior sampling.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration ( M ore) efficient reinforcement learning via posterior sampling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.961243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.458802Z digest=sha256:a1cf0c21b6db14a35b8d1a2fe3037ce6ae89f6ef774bd69cea93e8b26837c136

Observation edb6352d-6f2d-4a3e-beb1-dbcfbaeb60fa · outbound

This paper cites Deep exploration via bootstrapped DQN.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Deep exploration via bootstrapped DQN

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.951060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.462000Z digest=sha256:5c6833f6ba85248a590289b75ebb32fe3a4a572e0a665116bd5f6cc3b43634ab

Observation 11459e4e-5355-41db-9719-3281d0c2d48b · outbound

This paper cites J., Wen, Z., et al.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration J., Wen, Z., et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.940926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.465163Z digest=sha256:18932732e33409b9a533c4263e3d9a917ec5600704b45e69243cdc1030ad9cab

Observation 17e9035b-6bf4-4019-bcad-613d8f64e534 · outbound

This paper cites and Parr, R.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Parr, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.931009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.468243Z digest=sha256:8ce883cd8dc851cce175608646a380b4555ba919fda767b448a5d5cb5fe62f70

Observation 354e619f-3119-46f3-b542-e467fde92d3b · outbound

This paper cites Worst-case regret bounds for exploration via randomized value functions.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Worst-case regret bounds for exploration via randomized value functions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.920774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.471356Z digest=sha256:63bbf416b98e86ec41235716867a987fba2ed37b4038b87f16ea28e918b30b09

Observation 5ac2be83-181e-4420-8550-a36cf35dd0f5 · outbound

This paper cites and Van Roy, B.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Van Roy, B

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.474577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.474577Z digest=sha256:b9c7f71d32a8976565e2331bde9c572e64d12f7ad1c7337e9c084585c6516350

Observation 221eb884-702a-4c82-895b-e8399fb69849 · outbound

This paper cites Cooperative and Competitive Biases for Multi-Agent Reinforcement Learning.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Cooperative and Competitive Biases for Multi-Agent Reinforcement Learning

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:18:17.660714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.477793Z digest=sha256:1048ee06b2b8bbfe4348dc11e1ab66abe28db5aa5f0d1aa348d3e5a6e7c6c162

Observation 156f37e6-1807-42d9-a702-88bb985e3d3c · outbound

This paper cites and Leyton-Brown, K.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Leyton-Brown, K

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.904574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.481692Z digest=sha256:2896167400245bdd0a3664de9911b9ab11f76a2fcc48fba3f8090843e8d08eeb

Observation efffe54a-82a3-44c7-9f90-6ada333971f2 · outbound

This paper cites Multi-agent reinforcement learning: a critical survey.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Multi-agent reinforcement learning: a critical survey

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.894475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.484804Z digest=sha256:3c57c8294a51d7524c508e17bb0cc4ce0c6535726a21dfb6089041085724c5f4

Observation 7e3f5be5-f298-455e-97e7-d651ab6e6707 · outbound

This paper cites Concurrent reinforcement learning from customer interactions.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Concurrent reinforcement learning from customer interactions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.884630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.488356Z digest=sha256:5c6bbd59e91a83c68fa0f7f5698f6a6d292edf08d301bff46e765f0064e0aac0

Observation 0d8d3c58-3fb5-442a-910e-1fbcf410a061 · outbound

This paper cites AdaLead: A simple and robust adaptive greedy search algorithm for sequence design.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration AdaLead: A simple and robust adaptive greedy search algorithm for sequence design

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.491734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.491734Z digest=sha256:ef6c8a309c7db0c02f0e3cb0e004b3edc1c5b57d5d9a37900204a8b9a0c0c2b8

Observation 7b349bac-5711-47df-8239-923813f60ffc · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:18:17.874515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.495212Z digest=sha256:f64d4676b72ecc211f3935cd6d712b765c1d57c83834a59267a6722d5845ecb7

Observation d28d934f-8511-4603-ae20-c3589b6c9f0e · outbound

This paper cites an unresolved cited work.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.498464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.498464Z digest=sha256:85d6b166b4850f0a795b8010c7aed69cf18a3d75f177e9fd88c16dedd20f8bac

Observation 7b9ddedf-1641-4e2a-9273-88abf6a72103 · outbound

This paper cites and Littman, M.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Littman, M

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.858990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.501817Z digest=sha256:e37ce054ec21da328eb6f9b80d7bbf7d95d21926a072f4f362e45c793dc07d88

Observation bb8f331d-fd14-4160-a974-7e47818eecd8 · outbound

This paper cites A., Courville, A., and Bellemare, M.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration A., Courville, A., and Bellemare, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.849124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.505364Z digest=sha256:d9388fddd65d9e6f7f52a5c21a668d4f474ade6667e7174832774110bc837e71

Observation 05d7cf58-d422-40d2-bcf9-3bcfdc45efb3 · outbound

This paper cites Performance loss bounds for approximate value iteration with state aggregation.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Performance loss bounds for approximate value iteration with state aggregation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.839415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.508934Z digest=sha256:fbaf7605ec5a2f6f3e5860e86262740521d5b218a2f68cd6b1e2644c7ff4a94b

Observation 0b83ea6f-9e4b-46ea-b154-a90b8b0709e0 · outbound

This paper cites A concise introduction to multiagent systems and distributed artificial intelligence.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration A concise introduction to multiagent systems and distributed artificial intelligence

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.828359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.512242Z digest=sha256:63e4446d9879cb2c917302a129389c22e5a062b125f978ea8a26f3582c8f80a0

Observation afccec48-a659-4ca5-bd62-4b9d2bd4ca44 · outbound

This paper cites and Klabjan, D.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Klabjan, D

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.818329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.515848Z digest=sha256:03fea4064e620ee3384a552036340d6df36766017329f085b92958e085395787

Observation c062aba4-ce0e-438d-aaf1-960917a743f1 · outbound

This paper cites Multiagent systems: a modern approach to distributed artificial intelligence.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Multiagent systems: a modern approach to distributed artificial intelligence

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.807833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.519228Z digest=sha256:63698ddbf77ad5ff232c107d67852c55e0b8fce50a6f79234c28bb20771e5a2b

Observation 713f4f87-9a06-46a1-ad8e-3898b70d69c1 · outbound

This paper cites and Van Roy, B.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration and Van Roy, B

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.797372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.522440Z digest=sha256:7bda8a26724e159d0d58b3390c978b7dfa3ca848c671963e436bf2f99168e6f5

Observation 6da1f634-823e-4393-b905-5da185e79c4b · outbound

This paper cites Posterior sampling for continuing environments.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Posterior sampling for continuing environments

Reference 57

Resolution
verified exact
raw_fallback, observed 2026-08-10T16:18:17.635551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.525638Z digest=sha256:0d16ed40a6d6223f7446755facf489e6cc3d886ca7469c2d1d19f093fc673c9b

Observation 14a15f90-cae8-45d7-b3a3-688afba633b8 · outbound

This paper cites Frequentist regret bounds for randomized least-squares value iteration.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Frequentist regret bounds for randomized least-squares value iteration

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.788059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.528886Z digest=sha256:a762e19104fcaab023ab064635ea088c1d7274b52a28e5b440b45360ee3f4e98

Observation 1d45f6f4-e241-4dca-8371-45ce7b77c17c · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:17.532317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:18:17.532317Z digest=sha256:c7e70a141bfa9183e9ed43e877bc53fb5ddcc29376572f5319547f6d40298901

Observation 628b2879-55cb-4c71-90eb-6e91781539f8 · outbound

This paper cites Multi-agent cooperative reinforcement learning in 3d virtual world.

Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration Multi-agent cooperative reinforcement learning in 3d virtual world

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:18:17.772738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:18:17.535895Z digest=sha256:2bc73d32731296db1e09852b5e409cf51c9c71558076c30fa6537fc4c0355051

Pith citing papers

No inbound Pith citation observations are available.