Pith. sign in

Paper Citation Record · LEDGER

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

As of 19 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2502.02516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02516 v3

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:09.116234Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:54:00.193317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f81defeb-c1ed-4883-805f-fa8c0d194bce · outbound

This paper cites Define now the policy π(u|s) = P (u|s, π(s)).

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Define now the policy π(u|s) = P (u|s, π(s))

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.396885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.029865Z digest=sha256:ba7caaf2d487744ef4eed4b6a5d85eb0c81d3e9ab810f74b1337c9ea800b2ed8

Observation 18a61155-b45b-43a1-b567-3bf2bd923696 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.366261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.041873Z digest=sha256:653771e99dc4308b6f4448c52c84de39675390893b9e0ad36adb07ec202b9d14

Observation 095938ef-068c-4a9e-a1e7-1d1be90e62f5 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.381993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.035422Z digest=sha256:332018880bb2b7cfeb0f9259f341c3cd94d46aa43958d33aabc04089588bca2e

Observation 38ebd29f-8653-48fa-bda1-0ce07e72259f · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.269649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.091686Z digest=sha256:3988f32b01201cb28b3d8839b4ffd0ff5ffa0bb20a83b11420ad3f5f587ab7c3

Observation 9273a68d-b490-4b9a-985f-f1ef5d939d49 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.350101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.048682Z digest=sha256:e16b09aced638b362e7a5cd0025d7c88655a4af29bb1303742bc8a34f69e058b

Observation 5cf64e25-2ccc-4458-a23b-fa870086c120 · outbound

This paper cites sufficient.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation sufficient

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.335207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.054503Z digest=sha256:dd2cff1d5494e2d6af86baec3df57152c301168bbb01203677f9e80ce15600e8

Observation 9dbe6623-bd68-49ff-8fda-fbac3738d90c · outbound

This paper cites This choice encourages to select under-sampled actions for β >0, while for β = 0 we obtain a uniform forcing policy πf,t(a|s) = 1/A.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation This choice encourages to select under-sampled actions for β >0, while for β = 0 we obtain a uniform forcing policy πf,t(a|s) = 1/A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.319452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.064185Z digest=sha256:7fd18030d81c21d26bd269a7f17ee0f682292621141f79137b747283cbc69711

Observation 4c40abb8-cd42-4a39-ac6e-c57ced12dd03 · outbound

This paper cites Then, such solution induces an ergodic (irreducible and aperiodic) chain by Assumption 3.1 and Assumption 5.1.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Then, such solution induces an ergodic (irreducible and aperiodic) chain by Assumption 3.1 and Assumption 5.1

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.303737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.075723Z digest=sha256:e7f030c1973d0cb519572c12d556d30f0d14adc9e8c9ba62b6790f9e224cb608

Observation 9a96dabc-1749-4a6c-b267-a3e29c11f82d · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.287755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.084632Z digest=sha256:85d7d6231d5a9d4502a59f2c1ca329c2d1be124a9736c623d91b95c782b22492

Observation 9846c2a1-5e65-4537-9661-21aa528875c4 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.253084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.098369Z digest=sha256:b0bdee49535eb932dd476bc889a9d631383745534d6506bee6d3141906efa443

Observation dd763a3e-bbe3-4d47-a287-9b0accf62d59 · outbound

This paper cites On the other hand, in the single-policy scenario we use a default target policy policy πdef that is different for each environment.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation On the other hand, in the single-policy scenario we use a default target policy policy πdef that is different for each environment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.236354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.103954Z digest=sha256:6ab2ecc76a22eb0fb663dfd6929c88287c2585015565e1c099a513e2bd4d4cf5

Observation 73fd5463-4a96-43fc-8001-b8559bbc02da · outbound

This paper cites For the reward free case we use Rcanon to perform evaluation.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation For the reward free case we use Rcanon to perform evaluation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.216143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.110031Z digest=sha256:d150b52fb04b645b33a4246f8fc42149cf2956b00e25f507621845904b09b865

Observation 32fadcb6-8f4e-4849-9e20-e03dda2e8ab1 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.196812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.116234Z digest=sha256:5cb93f7689b2530c65373461866c02e58665d84457e296d6eb116df7cf2f4e91

Observation b03f0c00-833f-4b14-9a34-9ed35b420fcb · outbound

This paper cites Clustered KL-barycenter design for policy evaluation.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Clustered KL-barycenter design for policy evaluation

Reference 418

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:57:09.177406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T11:57:09.021310Z digest=sha256:ac421330f0a06aaa8db2f823fc632952a83e4c4d90ca5adc16f6bf97453c170d

Pith citing papers

Observation 661319d6-b19e-4270-8ce8-866f141e0195 · inbound

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning cites this paper.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.193317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.193317Z digest=sha256:9a119574768ac36a9c795991a468010606a64aa608b91809cad021c337b51dfa