Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning with Penalized Action Noise Injection

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02356 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:35.725134Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1da2545c-de84-4043-9793-2d18e1ffe030 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Offline Reinforcement Learning with Penalized Action Noise Injection Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.434048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.737091Z digest=sha256:2b393abed92f8222b92c1193c017760883069b89af4ade18d7b49c43e953eec4

Observation 3e5d8264-1aa1-4027-9faf-c087bdbf7e9a · outbound

This paper cites Layer Normalization.

Offline Reinforcement Learning with Penalized Action Noise Injection Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.842375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.842375Z digest=sha256:d86560c17b44fe8aad9deb2be119c6b1b33fc35748c49404332d6badc4a4a891

Observation dede0fb2-7726-4036-bc49-c3a9da87a142 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.961205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.961205Z digest=sha256:f755295a6593f756fcae603404ee39045f5ac130f8c9cd963ec5a43a2501e7f0

Observation f7d77ab3-7008-4262-a7e8-eb0341d5621b · outbound

This paper cites Score Regularized Policy Optimization through Diffusion Behavior.

Offline Reinforcement Learning with Penalized Action Noise Injection Score Regularized Policy Optimization through Diffusion Behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.067287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.067287Z digest=sha256:1c6893b3167d2500703c5f15df0d3aa4b0af379e1abecccd800458d332e99eac

Observation 30bc8fe0-d178-4f77-ac42-81ce3fe4e4d5 · outbound

This paper cites Diffusion Policies creating a Trust Region for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies creating a Trust Region for Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.141102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.141102Z digest=sha256:a0926c01c41c796f1c26c2c72ab8374a2d3789312ac104ffc6a7c49b39536108

Observation 270e3281-3349-405a-be63-20c824aff9ab · outbound

This paper cites Heavy-tailed denoising score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-tailed denoising score matching

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:38:36.017010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.216742Z digest=sha256:dceb46fbd3b60af261316e04a608d1f4401eeb9fafeeab0caa1b9166ec8c8866

Observation cfd8ec7d-c3f3-4b95-a8e1-9eed79c10ea8 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.343671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.343671Z digest=sha256:b758e98c9aee3f39ecef4b90f7dd480b2c31f5805e0228e190badaaca5dd1ef9

Observation 2458abd8-0a18-49dc-8f1e-74484648aefc · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection A minimalist approach to offline reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.427027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.427027Z digest=sha256:d0b13cf672440ddf29c7c9f3458e07582561208b1e8866436bc5f9be232a9b78

Observation fe4cf0f3-adb7-45c5-9de0-b0f5b52e3330 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Offline Reinforcement Learning with Penalized Action Noise Injection Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.494366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.494366Z digest=sha256:7e96f729aa8c5d80f24f6f6f444bd1c552f74cc1d3538211aad28280e691ba96

Observation 3b8f0a35-73b2-4bb6-91b2-d81f6d6e16cc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Off-policy deep reinforcement learning without exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.621106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.621106Z digest=sha256:341ea947d84600768e1de877de5f3f49114f478af2d3fb7cc81621b3d96f0762

Observation 0bdfee5c-e361-4565-82e3-156ebaf484f5 · outbound

This paper cites Calculus of variations.

Offline Reinforcement Learning with Penalized Action Noise Injection Calculus of variations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.743577Z digest=sha256:5841579cd5d04b6c6b44b070c919cb89769278f747faa029cf4f465258feb256

Observation 44a5caa0-2ef4-435c-bf9d-d4675b8f5884 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Offline Reinforcement Learning with Penalized Action Noise Injection IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.861799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.861799Z digest=sha256:e9d3a29f6fae78c6babe74b68c320943455c738af82bec2065b48faaf1054fbb

Observation 3fd83fcf-0547-4a0b-9dc1-ecabbe7bc4a1 · outbound

This paper cites Estimation of non-normalized statistical models by score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Estimation of non-normalized statistical models by score matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.978532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.978532Z digest=sha256:ddf3b27aa2c6200200f94fbab616fec5934338239530d893f6ea8758eebd8ed8

Observation 23d6c552-6feb-46af-97a5-0fbb323cb07f · outbound

This paper cites Understanding diffusion objectives as the elbo with simple data augmentation.

Offline Reinforcement Learning with Penalized Action Noise Injection Understanding diffusion objectives as the elbo with simple data augmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.232042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.074290Z digest=sha256:0bb00685671a6cb2dee40235db587b0638e53b31dcb233dee828601f3892a86c

Observation 0601121a-7f5e-46fb-ab8b-600e70a417b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Offline Reinforcement Learning with Penalized Action Noise Injection Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.201984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.201984Z digest=sha256:31dd7f62f86bddb162088080d8f60dabc6e0e7645c128fe67c331128fd5d012b

Observation 8dcf4a96-1f10-4f7d-8c6b-853a671b92f4 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning with Implicit Q-Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.288918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.288918Z digest=sha256:74505d464ce56c2d9173da6abc62fa0007a3d259f97f9d8f61da3992d184170d

Observation 09e7a1eb-a7a6-4c39-b6e4-c53cf57c1574 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Offline Reinforcement Learning with Penalized Action Noise Injection Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.405392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.405392Z digest=sha256:632f612f17565070e55e61af6f67b1db6139bf82b872e325b41c0fe2b7cff357

Observation bcf8ceda-bcaa-428f-bb39-bc5937405986 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Conservative q-learning for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.059986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.524554Z digest=sha256:b05f81b5d79aff7ffd8f83635f6648391c75b687a1f08f92b1cdb7ac5f6ba0ec

Observation 41b9557a-32c3-4a56-ba75-368c9b45e136 · outbound

This paper cites Reinforcement learning with augmented data.

Offline Reinforcement Learning with Penalized Action Noise Injection Reinforcement learning with augmented data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.879872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.627961Z digest=sha256:eb90889a953d8edc5a0f6a32b3b3f0d858d0dc4ce9c92adfa4693b9bef948eb0

Observation 55ec212d-70ea-418b-a4dd-af6133addaaa · outbound

This paper cites Batch reinforcement learning with hyperparameter gradients.

Offline Reinforcement Learning with Penalized Action Noise Injection Batch reinforcement learning with hyperparameter gradients

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.675965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.739155Z digest=sha256:417f58c8844f96d26170fa12383017a20af666ac6de7e3f5a45ec948de3e0363

Observation ec0ef95b-8d50-4a57-abaf-d3cabbd5cc84 · outbound

This paper cites Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.521833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.848855Z digest=sha256:ad42b45ace64586cfb03e4159046f643a9f960693460a9d645febc235ac01505

Observation d3ce1b0b-1a87-4a90-96ad-568a994fe870 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.258953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.985104Z digest=sha256:b352ff1f1f9b4c1807aae56096105d699171dd0370b21907a72aaebb0c95a16b

Observation 51d92faf-a390-4ba8-9e6e-a0c6d0d6d8af · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Offline Reinforcement Learning with Penalized Action Noise Injection AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.078884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.078884Z digest=sha256:39024bb98ce56cf423d10ace14267f6814772fb564f9bf66af37e66bffbd0222

Observation baf425d1-58da-4a6f-bb21-d574dbcd88a6 · outbound

This paper cites Anti-exploration by random network distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Anti-exploration by random network distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.079058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.197168Z digest=sha256:8134e748440ad4b4251820d592ae165244c49641582990a6c31f559bd2508c4d

Observation b65795e4-72c4-4282-bfeb-70a1979aaa5b · outbound

This paper cites Heavy-Tailed Diffusion Models.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-Tailed Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.311608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.311608Z digest=sha256:110a154f468c9ac2ea7ddc4f0a3a0236567731cc2ff9634c6a38c5fdc0ff9616

Observation 78e68a9c-3574-42a0-8af4-e66fd739dad3 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Offline Reinforcement Learning with Penalized Action Noise Injection DreamFusion: Text-to-3D using 2D Diffusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.429681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.429681Z digest=sha256:97309457df48a3eeba255d84386cb102cd96cf2a297169c4465e4bc4d14ff273

Observation c922708f-0235-4672-9a6f-680c9919918d · outbound

This paper cites Efficient differentiable simulation of articulated bodies.

Offline Reinforcement Learning with Penalized Action Noise Injection Efficient differentiable simulation of articulated bodies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.915029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.550267Z digest=sha256:f1362a1539fdf6501c0ca244c7b7ef1ce7e3f13adf6a20aa4b101d069ea41898

Observation 995cbc0a-a10c-42ba-8e3e-cbb70fa0c5ac · outbound

This paper cites Offline reinforcement learning as anti-exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline reinforcement learning as anti-exploration

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.743353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.671019Z digest=sha256:30da37378d179e35eeb807d8795de8e8360fcc81d7a1556aacc9685f92cbee36

Observation de5bffb6-547d-4f2c-a9f5-e3254fbdfef1 · outbound

This paper cites S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics.

Offline Reinforcement Learning with Penalized Action Noise Injection S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.568452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.772132Z digest=sha256:de1e46303f922ebcadf22e81d56a6dd729c1432d457da5bd29bd8d25a7d6a6e5

Observation bb38a4bf-2c58-4375-80ea-5a4aa65b7b1e · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Offline Reinforcement Learning with Penalized Action Noise Injection Generative modeling by estimating gradients of the data distribution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.862757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.862757Z digest=sha256:de4c9e74dde4ef8bb031e3013c7db79083bde2d938e2fa4c8d848da58c369a68

Observation 8a2cce95-1e29-4969-ab39-8b4baee2e0c2 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Offline Reinforcement Learning with Penalized Action Noise Injection Score-Based Generative Modeling through Stochastic Differential Equations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.981272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.981272Z digest=sha256:8c262bb13b5e524f64ebd649c22a328e32f719156de9a956742d31e6e116d59c

Observation af1d8a34-686b-4d5f-8006-e3459a95c306 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Revisiting the minimalist approach to offline reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.087356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.087356Z digest=sha256:490d0d1a8db720b972e3a46be6ab794bc1e7d425158fd2ad831c6ee38f0ca52c

Observation 78e6b11b-7441-4184-b72d-1864ae3d9c6f · outbound

This paper cites A connection between score matching and denoising autoencoders.

Offline Reinforcement Learning with Penalized Action Noise Injection A connection between score matching and denoising autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.208587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.208587Z digest=sha256:d73312b0228987b2e50a1d0c6d514da65a33d6cd80a9225a80a958fd74b24cf6

Observation 95dc4f19-86d8-4674-8e21-89a853191561 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.366704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.366704Z digest=sha256:c70f4394ba6ddd4e2de878679b4888c9ac899493971d1bd842fb0efd09bed4a8

Observation 0cfadee7-393b-42d9-8e70-a045f47d28c7 · outbound

This paper cites On scale mixtures of normal distributions.

Offline Reinforcement Learning with Penalized Action Noise Injection On scale mixtures of normal distributions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.366574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.479265Z digest=sha256:84f7f885565da53eeee80f8c00bbb35e2846cdc6ef17f58de49b9aff6e8f7937

Observation 993d3e30-6b2c-4ac3-b75d-ea23486c921c · outbound

This paper cites Exploration and Anti-Exploration with Distributional Random Network Distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Exploration and Anti-Exploration with Distributional Random Network Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.625180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.625180Z digest=sha256:c3d67c8932fa02ba9e4fcfbe07ffae426bf87da7a7cc196488007662dbdd73f4

Observation bea9b73e-a826-4506-ab75-db014269a869 · outbound

This paper cites Rorl: Robust offline reinforcement learning via conservative smoothing.

Offline Reinforcement Learning with Penalized Action Noise Injection Rorl: Robust offline reinforcement learning via conservative smoothing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.183894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.725134Z digest=sha256:c3f3d064b88d6043d573738c453570863b6df546a9ba95fb1534a70b2876d3d8

Pith citing papers

No inbound Pith citation observations are available.