Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning with Penalized Action Noise Injection

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02356 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:35.725134Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1da2545c-de84-4043-9793-2d18e1ffe030 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Offline Reinforcement Learning with Penalized Action Noise Injection Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.434048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.737091Z digest=sha256:f392404111cdaa60efab49c39f692468b4ac1f9ad06007f72b9c37be33c035d9

Observation 3e5d8264-1aa1-4027-9faf-c087bdbf7e9a · outbound

This paper cites Layer Normalization.

Offline Reinforcement Learning with Penalized Action Noise Injection Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.842375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.842375Z digest=sha256:cc48a0ac96942a6428ab55e67be838921acec72a459a9cb22d362edad17caf7a

Observation dede0fb2-7726-4036-bc49-c3a9da87a142 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.961205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.961205Z digest=sha256:933e7c2ba52362522b530a730dbd9483ae8c053acbb06569f2fd13112d9b1efd

Observation f7d77ab3-7008-4262-a7e8-eb0341d5621b · outbound

This paper cites Score Regularized Policy Optimization through Diffusion Behavior.

Offline Reinforcement Learning with Penalized Action Noise Injection Score Regularized Policy Optimization through Diffusion Behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.067287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.067287Z digest=sha256:cb2f6a791985caef6ac2baf836bfb20ebb9e4922d3c48d948d43ac446588f70d

Observation 30bc8fe0-d178-4f77-ac42-81ce3fe4e4d5 · outbound

This paper cites Diffusion Policies creating a Trust Region for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies creating a Trust Region for Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.141102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.141102Z digest=sha256:1b70b46a5ddc93617bea695d9d8aa3a1de9e48a62de72ea19e2f958ba9f64720

Observation 270e3281-3349-405a-be63-20c824aff9ab · outbound

This paper cites Heavy-tailed denoising score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-tailed denoising score matching

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:38:36.017010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.216742Z digest=sha256:62013020fd568f76896c0597b9166c9fbb105016af462b4a5099e83e52a930fd

Observation cfd8ec7d-c3f3-4b95-a8e1-9eed79c10ea8 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.343671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.343671Z digest=sha256:10123c65941473cb3d8bc6058338a7b9a04483adfdf06b6ea9e90910ef3efa18

Observation 2458abd8-0a18-49dc-8f1e-74484648aefc · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection A minimalist approach to offline reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.427027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.427027Z digest=sha256:c37edd34a49e2dddaf9618077cc23cb0dc3ba483eabc78a0d52f302ab92a53d9

Observation fe4cf0f3-adb7-45c5-9de0-b0f5b52e3330 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Offline Reinforcement Learning with Penalized Action Noise Injection Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.494366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.494366Z digest=sha256:4ff4415ae9d9d6c2129776b9c20c85bb0c0e2d28b26f86380ce2ae5064881c5b

Observation 3b8f0a35-73b2-4bb6-91b2-d81f6d6e16cc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Off-policy deep reinforcement learning without exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.621106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.621106Z digest=sha256:0dacf84f8c76733f18917ccc07a70e81bf39fd2b181421bc647313aaa3efc830

Observation 0bdfee5c-e361-4565-82e3-156ebaf484f5 · outbound

This paper cites Calculus of variations.

Offline Reinforcement Learning with Penalized Action Noise Injection Calculus of variations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.743577Z digest=sha256:4404dc377dcf05e46bbae5895d912a9566c8c76a7e5dda1770389423bfcd3d7f

Observation 44a5caa0-2ef4-435c-bf9d-d4675b8f5884 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Offline Reinforcement Learning with Penalized Action Noise Injection IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.861799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.861799Z digest=sha256:0910d43b4f82e29214e7462f6a35e4ec4d6e44c510bca945f3f2c032c3c9a4ab

Observation 3fd83fcf-0547-4a0b-9dc1-ecabbe7bc4a1 · outbound

This paper cites Estimation of non-normalized statistical models by score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Estimation of non-normalized statistical models by score matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.978532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.978532Z digest=sha256:2fbacbc0f3e15f5c53797f60300af16f1d8258012e250297440a2fe062ac2307

Observation 23d6c552-6feb-46af-97a5-0fbb323cb07f · outbound

This paper cites Understanding diffusion objectives as the elbo with simple data augmentation.

Offline Reinforcement Learning with Penalized Action Noise Injection Understanding diffusion objectives as the elbo with simple data augmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.232042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.074290Z digest=sha256:6c0e142d02ff46d19eed4b925a49afab489b71382fa3a0a8d3477cfd7f78da9e

Observation 0601121a-7f5e-46fb-ab8b-600e70a417b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Offline Reinforcement Learning with Penalized Action Noise Injection Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.201984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.201984Z digest=sha256:4b8beed47bcbc6299827906c1011e90138acd314b0a8ccffb86b8fc62fddb2d9

Observation 8dcf4a96-1f10-4f7d-8c6b-853a671b92f4 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning with Implicit Q-Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.288918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.288918Z digest=sha256:5c3aeb38420c1e31b3119835b6a335c8cbd9b3cdfee67826c172a80b4bddcd3f

Observation 09e7a1eb-a7a6-4c39-b6e4-c53cf57c1574 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Offline Reinforcement Learning with Penalized Action Noise Injection Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.405392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.405392Z digest=sha256:7c4c26c3a0da480cdbf92ea66b3146f9b4a475f77ced8c3f6c45c7bd411dde70

Observation bcf8ceda-bcaa-428f-bb39-bc5937405986 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Conservative q-learning for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.059986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.524554Z digest=sha256:73a998051b3eb11891c643368ea9563cd6406b73bb66929763225d2825c806b7

Observation 41b9557a-32c3-4a56-ba75-368c9b45e136 · outbound

This paper cites Reinforcement learning with augmented data.

Offline Reinforcement Learning with Penalized Action Noise Injection Reinforcement learning with augmented data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.879872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.627961Z digest=sha256:6d248a21ae20f9b303cb38998831f0f2bb12269bfa62822df93767e674554250

Observation 55ec212d-70ea-418b-a4dd-af6133addaaa · outbound

This paper cites Batch reinforcement learning with hyperparameter gradients.

Offline Reinforcement Learning with Penalized Action Noise Injection Batch reinforcement learning with hyperparameter gradients

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.675965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.739155Z digest=sha256:fd9fe0095a8ae45dd885255d43e35eddd4aca7ef0a5ad437231817bd7632c486

Observation ec0ef95b-8d50-4a57-abaf-d3cabbd5cc84 · outbound

This paper cites Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.521833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.848855Z digest=sha256:ce992636886781fe865a9860e6c4c1478bc38735951a7e2dde121397c8060bd8

Observation d3ce1b0b-1a87-4a90-96ad-568a994fe870 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.258953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.985104Z digest=sha256:5c241cd84ccfc3696283da75a6dedd84f35f4e5240f27680f9f9e4b29a4fbd7e

Observation 51d92faf-a390-4ba8-9e6e-a0c6d0d6d8af · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Offline Reinforcement Learning with Penalized Action Noise Injection AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.078884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.078884Z digest=sha256:7c8a8244e79a21c3ee93c0e2df104cbd2474ee3f88f35b50be7d70416af11509

Observation baf425d1-58da-4a6f-bb21-d574dbcd88a6 · outbound

This paper cites Anti-exploration by random network distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Anti-exploration by random network distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.079058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.197168Z digest=sha256:c1719272bf1cfc679afaa0029811726b0bb5e4c66484539e9ffa7bf07c8de6b6

Observation b65795e4-72c4-4282-bfeb-70a1979aaa5b · outbound

This paper cites Heavy-Tailed Diffusion Models.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-Tailed Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.311608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.311608Z digest=sha256:11722bcbcedd9abcc5998a6a699161f0a64267ceeb9f3c4f34d90693bd6082be

Observation 78e68a9c-3574-42a0-8af4-e66fd739dad3 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Offline Reinforcement Learning with Penalized Action Noise Injection DreamFusion: Text-to-3D using 2D Diffusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.429681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.429681Z digest=sha256:93bacbb188bfd207e04700ec3d493abbae37263634a5a94c36415cd78b9b69e6

Observation c922708f-0235-4672-9a6f-680c9919918d · outbound

This paper cites Efficient differentiable simulation of articulated bodies.

Offline Reinforcement Learning with Penalized Action Noise Injection Efficient differentiable simulation of articulated bodies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.915029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.550267Z digest=sha256:3a9597e7dd723b790d0eba581d6ee385afcf182b11b98b37402537cf9f2a8552

Observation 995cbc0a-a10c-42ba-8e3e-cbb70fa0c5ac · outbound

This paper cites Offline reinforcement learning as anti-exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline reinforcement learning as anti-exploration

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.743353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.671019Z digest=sha256:6fac41bfc7b6b8eb84575610313b51480bf21730bc139036c512c9e2b73efa20

Observation de5bffb6-547d-4f2c-a9f5-e3254fbdfef1 · outbound

This paper cites S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics.

Offline Reinforcement Learning with Penalized Action Noise Injection S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.568452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.772132Z digest=sha256:e166d2acbaad1e8c42259a1b35e85d7db954b94403998adb21fadb1e16b59a32

Observation bb38a4bf-2c58-4375-80ea-5a4aa65b7b1e · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Offline Reinforcement Learning with Penalized Action Noise Injection Generative modeling by estimating gradients of the data distribution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.862757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.862757Z digest=sha256:23d5d91f86b42c28c910cf779589ff75adb439f66853312b4b947812b6f32ea5

Observation 8a2cce95-1e29-4969-ab39-8b4baee2e0c2 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Offline Reinforcement Learning with Penalized Action Noise Injection Score-Based Generative Modeling through Stochastic Differential Equations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.981272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.981272Z digest=sha256:5a5933795686906f102e96abef0671dc7525e12d41dd5b9a24536468ca530c63

Observation af1d8a34-686b-4d5f-8006-e3459a95c306 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Revisiting the minimalist approach to offline reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.087356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.087356Z digest=sha256:26f25bbfc195c1e6afe53e636591b78fd9d55503ab98d6fa9f59d0910f4ed9cd

Observation 78e6b11b-7441-4184-b72d-1864ae3d9c6f · outbound

This paper cites A connection between score matching and denoising autoencoders.

Offline Reinforcement Learning with Penalized Action Noise Injection A connection between score matching and denoising autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.208587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.208587Z digest=sha256:31cae1119d44079b5c80e93800e9cb55cac4c1bd685b189c82e921a5c734c4e2

Observation 95dc4f19-86d8-4674-8e21-89a853191561 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.366704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.366704Z digest=sha256:cc5a9ca045c3087592ddbe074d71584eac6f100bd98263a99e3f6c9bab6f9c9d

Observation 0cfadee7-393b-42d9-8e70-a045f47d28c7 · outbound

This paper cites On scale mixtures of normal distributions.

Offline Reinforcement Learning with Penalized Action Noise Injection On scale mixtures of normal distributions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.366574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.479265Z digest=sha256:e5b017c3bb78b34d78101649af140442749f4604aebbe5501c8c5a03f1cb6510

Observation 993d3e30-6b2c-4ac3-b75d-ea23486c921c · outbound

This paper cites Exploration and Anti-Exploration with Distributional Random Network Distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Exploration and Anti-Exploration with Distributional Random Network Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.625180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.625180Z digest=sha256:3ba80ee4990b42bf2c2fadc5dec0e43ad4d5b7417791e60c5c3c19dc4267319e

Observation bea9b73e-a826-4506-ab75-db014269a869 · outbound

This paper cites Rorl: Robust offline reinforcement learning via conservative smoothing.

Offline Reinforcement Learning with Penalized Action Noise Injection Rorl: Robust offline reinforcement learning via conservative smoothing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.183894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.725134Z digest=sha256:7bb9e37bf9d91dffafffba7dc762faa693b7e798ef40b5e06a31953e18788e27

Pith citing papers

No inbound Pith citation observations are available.