Pith. sign in

Paper Citation Record · LEDGER

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.21841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21841 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:51.668145Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 505562c8-8319-48b4-a76b-f495dc7b1a99 · outbound

This paper cites write newline.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:28.970873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:34:28.970873Z digest=sha256:1fd1af9a4694abe72032fe5d2f07360597677d5bdc5742a3e00f41c93bd958ab

Observation 630fe61f-14f5-4025-aa51-0741d651fe4b · outbound

This paper cites Constrained policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.539697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.027164Z digest=sha256:c0e27bc9ceaf6b4aa0ba72dc76ba115cc1073fc5cd39fc0b3f07a3ea991f0dbb

Observation 00939444-31d3-4e08-bcf8-9f10515a817e · outbound

This paper cites Constrained Markov decision processes, volume 7.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.227495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.159068Z digest=sha256:effd9533215dde15916be954539d525399beb0c40acb48c738780a2537e5a540

Observation 702aad4e-aa41-47b2-bc61-d37059c9958e · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.051680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.450316Z digest=sha256:f93e8cb24f5181250560bff2016c8624873f4dd8ef152fa69b35aaf6fb3a7c00

Observation f732cb5a-352a-4cee-a1b6-46c437e7cb63 · outbound

This paper cites G., Osband, I., and Munos, R.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints G., Osband, I., and Munos, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:47.683128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:47.683128Z digest=sha256:a651fbec9807a92dff6f24dd9c32636420a10a8ac0c16259ff5e5e16fcfe0c7e

Observation 11fef369-1567-4313-a78b-c6b810038091 · outbound

This paper cites S., Agarwal, M., Koppel, A., and Aggarwal, V.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints S., Agarwal, M., Koppel, A., and Aggarwal, V

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.781747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.767912Z digest=sha256:eae4a1020b38ab4ad9f460471b4e162a61da888faccd0f0cfb4365900859ceaa

Observation eff5fcfe-0b06-4347-bcf2-9688eb8b64b2 · outbound

This paper cites DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.527372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.885604Z digest=sha256:f4b651f07aa9242bc3acf2c80d6b8ea25be0bdd6b310ca917c9346dceecfea77

Observation 3347a759-18a2-4458-a690-11666c403118 · outbound

This paper cites Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.355079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.002331Z digest=sha256:cba2e162103368e4590e8577f365f705a01e3f513b6136e1a1c813d5642634b5

Observation 900a97dd-cfee-456f-953d-bbc0c71787bc · outbound

This paper cites Learning infinite-horizon average-reward markov decision process with constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning infinite-horizon average-reward markov decision process with constraints

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.621924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.098233Z digest=sha256:bbfcaca6b3fdfae942950f33e101779167005001f8d5a6178b6b8d5686d3bf61

Observation 06e5816d-1349-4dd4-b8ec-d67ed52b0e08 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Risk-constrained reinforcement learning with percentile risk criteria

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.335859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.234339Z digest=sha256:2b049092ff02f4c21325b1bea0b9201781b796843e2a9b70042cb4b498e245a8

Observation 8fcd96a9-0e4c-43af-b02f-be102f098bb2 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.018850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.336113Z digest=sha256:04f407f1496a47763d8244456d5fc80b10b6278126ed11031080407842a3b92b

Observation cdc0c94a-c330-4168-81ad-7bdc10a34556 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient safe exploration via primal-dual policy optimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.831806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.467808Z digest=sha256:f1f5df1e487bd01791e7c7eb128c6840cc8d1225ec1273955e4e1d4cfb079574

Observation 6e50368c-c30d-4027-b003-6cdfbcd64dc7 · outbound

This paper cites Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.098564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.569789Z digest=sha256:38bfd0e6305c53158c5a1d386f91037216b3d7eecd536b15ac4ca3019387ccfa

Observation 9b1fab1b-2b67-43ce-a413-b9e645253a4c · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Exploration-Exploitation in Constrained MDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:48.684747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:48.684747Z digest=sha256:9f26ca9a41639f8224fda89adc68885eddb2ea8ffe2e09598a2e233c6cc2af02

Observation 3ad0cfde-e1bf-4a66-b104-c999d0353657 · outbound

This paper cites A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:52.838965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.792791Z digest=sha256:a508327db5a7cacc6c8c8604a0c292827b42029c2df841c20f505dc7e2dad95d

Observation e064fe05-10de-4988-a9a1-0ebb44156973 · outbound

This paper cites Provably efficient model-free constrained rl with linear function approximation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free constrained rl with linear function approximation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.615244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.894556Z digest=sha256:0facd79bf22b51e561501e135bef6c1cdc560b9c0a8182052e5b02e1dbd16e7b

Observation d06a5db2-2c72-4b9a-bcbc-5682e2b3ec77 · outbound

This paper cites Online convex optimization with hard constraints: Towards the best of two worlds and beyond.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Online convex optimization with hard constraints: Towards the best of two worlds and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.433971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.064830Z digest=sha256:60c42d3539ce2957dcd523ec3f1f5966515d72ff698a99b6eea691b39d6ff1f9

Observation 55b6921b-dfe2-4b24-92c6-c7e51f6ea4aa · outbound

This paper cites Safe reinforcement learning on autonomous vehicles.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Safe reinforcement learning on autonomous vehicles

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.252065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.160323Z digest=sha256:4ce26b5fea9adb15b6fc15966f46fd4b372dae0d26375c933586338ac681b4be

Observation cd4ebc55-86eb-4332-9128-e09e759b68b5 · outbound

This paper cites Learning adversarial markov decision processes with bandit feedback and unknown transition.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning adversarial markov decision processes with bandit feedback and unknown transition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.073085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.233974Z digest=sha256:535ed688a47e92248569a520e0b419ff7830614a7bb6f8edf5a9c87a83561b8c

Observation 4cc7faf8-aad4-4a61-8a87-fc00818c6150 · outbound

This paper cites Asymptotically Optimal Information-Directed Sampling.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Asymptotically Optimal Information-Directed Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.598973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.315365Z digest=sha256:ea904d2d84c0be3450b8fbfd1918ea794f7b76974d1705e104623250b8893f08

Observation 3152b411-b341-40c6-8c39-757ffd025609 · outbound

This paper cites A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.425405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.425405Z digest=sha256:5caf9b61ed4a804f9f97f88f03d38a169115015360d6efab86de4e4763317e38

Observation a8af659f-5cf7-463d-b447-43f7d5501947 · outbound

This paper cites An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.517407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.517407Z digest=sha256:7ef6e5a33163c3f8d05478ffd3d1f8d2068595c17c728381d6a933095e149e7d

Observation 0ae5c9b7-2a55-4a3d-8f5c-2bddaf7c3ce1 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained MDPs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.912890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.610537Z digest=sha256:77371aacbd9e54ae7c53a530d85090c5664a0685294ce8559d165ee7dde22e9e

Observation 43ca9dfd-247a-45e0-8f8d-a91c9e1a2c08 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained mdps.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained mdps

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.718689Z digest=sha256:1587f94d645d1569d1c16089147b0ebf1a3f1d25cc0bdfb75bb39cc67400aecd

Observation 30023b5d-0b14-481a-8259-27aae1a15833 · outbound

This paper cites Policy optimization in adversarial mdps: Improved exploration via dilated bonuses.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Policy optimization in adversarial mdps: Improved exploration via dilated bonuses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.501758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.844197Z digest=sha256:565d0d9d6216bc2db6b818be8216490a779af06c07ad2aecf82278c4d1f0d73b

Observation b0a2135b-3cbe-424a-981b-d9ac16ab7d0a · outbound

This paper cites Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.965851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.965851Z digest=sha256:03f0218033961f366fc9aa178c903b6783582af4114c1e2b86b838606368e54d

Observation 132bc39d-c3d1-47da-bf21-fd3bedf6e30a · outbound

This paper cites Truly No-Regret Learning in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Truly No-Regret Learning in Constrained MDPs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.106454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:50.106454Z digest=sha256:6df44d82109e52dfa7b2705c414ddf37403fa8ef14689c6d9271b0a60f0118ff

Observation 4bd08f26-9eb0-45c3-84b4-49c23f9636e1 · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:55.253773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.236838Z digest=sha256:5436cc1b8888e5a31379e8db74572f5b4f46be8e5227553cfb0e86166402cd8b

Observation d496fcf0-1dd2-49f9-b471-fda868e43a9b · outbound

This paper cites Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.038809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.374739Z digest=sha256:1c721a51765c91766847790269baa9506c751004898877df5eb32fe5e16c0d2a

Observation 1d7336e1-2e04-4a89-9b5c-f897fe1b595d · outbound

This paper cites and Sridharan, K.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Sridharan, K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.836763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.535752Z digest=sha256:338b9f9632c9ae271a6c36aeb0647e28843ff6eda1d19c71d8509747c42bb653

Observation 5435887d-aee8-47e5-be4a-1626685ea8d8 · outbound

This paper cites Learning in Markov Decision Processes under Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning in Markov Decision Processes under Constraints

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.311657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.654786Z digest=sha256:40a7a4cca28833e929261bbaaa1c4adbd4fdb02fb740c5e7e57b95492dfba9ca

Observation 69f3302e-3541-4a29-b121-6ab7b2a4be2f · outbound

This paper cites Optimal Algorithms for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Algorithms for Online Convex Optimization with Adversarial Constraints

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.145545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.768281Z digest=sha256:88e8b2a08c0e013b1dad2b56610c5568a8d4a252855c63f1ae53d7eafcbee2b6

Observation a182d344-fcc4-42ee-bf95-7cbc40d03d95 · outbound

This paper cites Learning Adversarial MDPs with Stochastic Hard Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning Adversarial MDPs with Stochastic Hard Constraints

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.044652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:51.044652Z digest=sha256:9b308088e23c60ff7494beb1f4504df9d65d33bc4f75902074ab3ff4d78385e3

Observation b8adca20-6eb5-4929-a697-dde12424916b · outbound

This paper cites Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:51.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.152297Z digest=sha256:fdce9971b47c3f8adf82efb25bb417134be15255bea63c3871e73786f1a8fe0d

Observation 18847069-32b1-4364-aee1-77751cd6bdd9 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.549056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.294799Z digest=sha256:ab199337a0f925417a6d913fdb99253adba60506209a6c2b85148cdbd4fdefc4

Observation ac158d50-b7f6-4f0d-928e-2102d5c876fb · outbound

This paper cites A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.348961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.366080Z digest=sha256:df768826ad53265375d28a0b3986f35cfc9462ef8d0ca32feca39a2f93ca114c

Observation 10b44870-fe67-4b09-aa82-f812732727e1 · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDP s.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free algorithms for non-stationary CMDP s

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.171541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.499356Z digest=sha256:c377feecba85c6e090cb6895d8ad4791224efa371c8cd0e551d12652b71f75f4

Observation 531160d9-5088-45ee-a051-a9a4dcc157ab · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:53.960811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.570309Z digest=sha256:8bff2be530c09fecb9fd2eb5132cf11d129ffa98d2c88c7973dc87de1aa36828

Observation 42c9f1b7-449e-4905-b842-e43a45cd2e29 · outbound

This paper cites and Ugot, O.-A.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Ugot, O.-A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:53.728331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.668145Z digest=sha256:6af904f4a4eb61a727cce1ce0db61ee0e087513a32afdd4de77d91b78067b2bd

Pith citing papers

No inbound Pith citation observations are available.