Pith. sign in

Paper Citation Record · LEDGER

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

As of 15 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.21841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21841 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:51.668145Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 505562c8-8319-48b4-a76b-f495dc7b1a99 · outbound

This paper cites write newline.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:28.970873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:34:28.970873Z digest=sha256:8f8a24d8817a37938b3e13a55b49364ddc81357a49e75b1bdd2220cf74355ef9

Observation 630fe61f-14f5-4025-aa51-0741d651fe4b · outbound

This paper cites Constrained policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.539697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.027164Z digest=sha256:8488daa501b26f2f510114a4ceb3a3d7c5637e9c94eea04eaa665b8f15321c57

Observation 00939444-31d3-4e08-bcf8-9f10515a817e · outbound

This paper cites Constrained Markov decision processes, volume 7.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.227495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.159068Z digest=sha256:8a27f10e0fa6883d44e7398a92b5a84659873d5a1525f23f503991d5cf970ec8

Observation 702aad4e-aa41-47b2-bc61-d37059c9958e · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.051680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.450316Z digest=sha256:92f591463a55a888202b4d234d192e7b1896123b410b3cc5d3549615846be880

Observation f732cb5a-352a-4cee-a1b6-46c437e7cb63 · outbound

This paper cites G., Osband, I., and Munos, R.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints G., Osband, I., and Munos, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:47.683128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:47.683128Z digest=sha256:f31bee78004abdd9377c4486ac12f377e5f73f17ec05ae9d0b31b4f523b6970a

Observation 11fef369-1567-4313-a78b-c6b810038091 · outbound

This paper cites S., Agarwal, M., Koppel, A., and Aggarwal, V.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints S., Agarwal, M., Koppel, A., and Aggarwal, V

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.781747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.767912Z digest=sha256:5e44a782f28afe4414462877dbd1382315d8b4ba69c0e2300a73a20dfcaf1c00

Observation eff5fcfe-0b06-4347-bcf2-9688eb8b64b2 · outbound

This paper cites DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.527372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.885604Z digest=sha256:66e340460ead0e4ac296b3c03b70d0ce435e19a8bc391eea021008d29fc8a07b

Observation 3347a759-18a2-4458-a690-11666c403118 · outbound

This paper cites Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.355079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.002331Z digest=sha256:6a38caf1ec938efd75b0438ef5216d54ae8477c7f66b355a0f8c027640542682

Observation 900a97dd-cfee-456f-953d-bbc0c71787bc · outbound

This paper cites Learning infinite-horizon average-reward markov decision process with constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning infinite-horizon average-reward markov decision process with constraints

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.621924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.098233Z digest=sha256:d43e13d3fc895c26fe257fb9a34837922c82c23c000d278c17d29265d830eafd

Observation 06e5816d-1349-4dd4-b8ec-d67ed52b0e08 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Risk-constrained reinforcement learning with percentile risk criteria

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.335859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.234339Z digest=sha256:3a19ddc1e90e335f32ccaa9e7426f7c0f12ec605ec76e459b5d7e4ba72c7e65f

Observation 8fcd96a9-0e4c-43af-b02f-be102f098bb2 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.018850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.336113Z digest=sha256:a614d6f6660509079693a3c432b0a676f45730b8b40d1571326a7e1efa069743

Observation cdc0c94a-c330-4168-81ad-7bdc10a34556 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient safe exploration via primal-dual policy optimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.831806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.467808Z digest=sha256:d802008fed9c2ca9ddfcd71233b11fd32de5d58735abb9616e9c8c0198940fd6

Observation 6e50368c-c30d-4027-b003-6cdfbcd64dc7 · outbound

This paper cites Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.098564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.569789Z digest=sha256:380cc526dc3a9eb80ecaf3fb9b877c348fb9a8feda82dcbe58a76cc16e1af9ce

Observation 9b1fab1b-2b67-43ce-a413-b9e645253a4c · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Exploration-Exploitation in Constrained MDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:48.684747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:48.684747Z digest=sha256:b5440d0c45296680cbe9f8810c0d8d2216ac94a482adfd7c037b2b3046c8706d

Observation 3ad0cfde-e1bf-4a66-b104-c999d0353657 · outbound

This paper cites A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:52.838965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.792791Z digest=sha256:78d3e60b6d8e254839c8825fa10b9583f225fe3c9069ec47df5b1123f7504091

Observation e064fe05-10de-4988-a9a1-0ebb44156973 · outbound

This paper cites Provably efficient model-free constrained rl with linear function approximation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free constrained rl with linear function approximation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.615244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.894556Z digest=sha256:8e2cd77198a9f68bf7b0bbed5be55dd4fc0b2133b8bc880bf8097401baa99b21

Observation d06a5db2-2c72-4b9a-bcbc-5682e2b3ec77 · outbound

This paper cites Online convex optimization with hard constraints: Towards the best of two worlds and beyond.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Online convex optimization with hard constraints: Towards the best of two worlds and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.433971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.064830Z digest=sha256:f09bc64c270293763bbfdf7929339bfae5b9bc50d812a5cb42c9ce686533b829

Observation 55b6921b-dfe2-4b24-92c6-c7e51f6ea4aa · outbound

This paper cites Safe reinforcement learning on autonomous vehicles.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Safe reinforcement learning on autonomous vehicles

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.252065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.160323Z digest=sha256:412ec7c9b44902a5f65e0fa4d94457209ab92135bb8afb9857d959ac0a284730

Observation cd4ebc55-86eb-4332-9128-e09e759b68b5 · outbound

This paper cites Learning adversarial markov decision processes with bandit feedback and unknown transition.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning adversarial markov decision processes with bandit feedback and unknown transition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.073085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.233974Z digest=sha256:dac86d886e30daa8f637c6fa80ee31410182f74cdae892aeac11aad423fbfd59

Observation 4cc7faf8-aad4-4a61-8a87-fc00818c6150 · outbound

This paper cites Asymptotically Optimal Information-Directed Sampling.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Asymptotically Optimal Information-Directed Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.598973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.315365Z digest=sha256:83218bee4e51f420308edf9326160c3309ddf141bbd05e1297b62b48ffc4bdee

Observation 3152b411-b341-40c6-8c39-757ffd025609 · outbound

This paper cites A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.425405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.425405Z digest=sha256:ee508233fbafb50f266ee70909c81680ecd5508035a8b5e4e7a72cf724d70c4c

Observation a8af659f-5cf7-463d-b447-43f7d5501947 · outbound

This paper cites An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.517407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.517407Z digest=sha256:241a8bf44ad700257fa293e9a438f0d6c41916531c1d623bba667e6109606a6d

Observation 0ae5c9b7-2a55-4a3d-8f5c-2bddaf7c3ce1 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained MDPs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.912890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.610537Z digest=sha256:f44badd447d83841f38484cc2cfa5eedb8efd38a46053fe79a287fac2beb3917

Observation 43ca9dfd-247a-45e0-8f8d-a91c9e1a2c08 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained mdps.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained mdps

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.718689Z digest=sha256:2155973966b164109cb0368514532ec8822db0ce6cecbd3c7db4932cf1a85d43

Observation 30023b5d-0b14-481a-8259-27aae1a15833 · outbound

This paper cites Policy optimization in adversarial mdps: Improved exploration via dilated bonuses.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Policy optimization in adversarial mdps: Improved exploration via dilated bonuses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.501758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.844197Z digest=sha256:7de88f6e16663c2b9c7986756bb418804326e89ae3302a8266cbd2fecdb40768

Observation b0a2135b-3cbe-424a-981b-d9ac16ab7d0a · outbound

This paper cites Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.965851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.965851Z digest=sha256:d6e52dd275ea3d4a833c7677b1731b1db2d3b2c88c03606b3ccefea082d95c09

Observation 132bc39d-c3d1-47da-bf21-fd3bedf6e30a · outbound

This paper cites Truly No-Regret Learning in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Truly No-Regret Learning in Constrained MDPs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.106454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:50.106454Z digest=sha256:37bccc3d53da389f76f98974362874e7e268e00e7543a586a28706886cf232a7

Observation 4bd08f26-9eb0-45c3-84b4-49c23f9636e1 · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:55.253773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.236838Z digest=sha256:822f92a6e018ca5dfcea9cfe1b420d6bcc152a2b2859fb842c79b4b2957697cb

Observation d496fcf0-1dd2-49f9-b471-fda868e43a9b · outbound

This paper cites Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.038809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.374739Z digest=sha256:265a95a1808d9491dd90e8ecd607f58a5090ff72f6867307750426a658ac9280

Observation 1d7336e1-2e04-4a89-9b5c-f897fe1b595d · outbound

This paper cites and Sridharan, K.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Sridharan, K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.836763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.535752Z digest=sha256:f2745aaa4e67551bdbb5a98aea562545e3cd297d7cb378936dc0bacf09162975

Observation 5435887d-aee8-47e5-be4a-1626685ea8d8 · outbound

This paper cites Learning in Markov Decision Processes under Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning in Markov Decision Processes under Constraints

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.311657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.654786Z digest=sha256:251cc29a669f762e8041df45e39792f7f3838be45a09fc7984b7e1db21120541

Observation 69f3302e-3541-4a29-b121-6ab7b2a4be2f · outbound

This paper cites Optimal Algorithms for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Algorithms for Online Convex Optimization with Adversarial Constraints

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.145545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.768281Z digest=sha256:4fc2ffc5f1f77611b3f16f79fa9e399c60034229ca8b7ca9d1c69ce097d75ec2

Observation a182d344-fcc4-42ee-bf95-7cbc40d03d95 · outbound

This paper cites Learning Adversarial MDPs with Stochastic Hard Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning Adversarial MDPs with Stochastic Hard Constraints

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.044652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:51.044652Z digest=sha256:d3b63f007e52ed3c1a0a0efbb7b74134154be4d90d2deb03b66e29536d8ef02a

Observation b8adca20-6eb5-4929-a697-dde12424916b · outbound

This paper cites Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:51.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.152297Z digest=sha256:6beed6bd85197c2245b223eab95f304e9d5ebaa4739cdeffa5694b0045b59874

Observation 18847069-32b1-4364-aee1-77751cd6bdd9 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.549056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.294799Z digest=sha256:558d9082b30a14fd436b49f898c663672b4caf341b6dc0179ddc35e523c0c5ec

Observation ac158d50-b7f6-4f0d-928e-2102d5c876fb · outbound

This paper cites A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.348961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.366080Z digest=sha256:5a16e0dda35aacf6ae7ab5e2fe2324decf7493e6d309a0695e11719d0fe74189

Observation 10b44870-fe67-4b09-aa82-f812732727e1 · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDP s.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free algorithms for non-stationary CMDP s

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.171541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.499356Z digest=sha256:297386f9281b6d2d0c0cb1097e669027893cfd7a25c312611d635c20dbbe4669

Observation 531160d9-5088-45ee-a051-a9a4dcc157ab · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:53.960811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.570309Z digest=sha256:85922181547246dfc9a361cf42208442b4e695a6094615322b88b07559526684

Observation 42c9f1b7-449e-4905-b842-e43a45cd2e29 · outbound

This paper cites and Ugot, O.-A.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Ugot, O.-A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:53.728331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.668145Z digest=sha256:306c09e83420f9ee83457d6d14e7fc1a1a7a57c550d7982ba679676c1113366b

Pith citing papers

No inbound Pith citation observations are available.