Pith. sign in

Paper Citation Record · LEDGER

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2506.14058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14058 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:05:34.935087Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T20:52:08.062450Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T20:54:21.647943Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2deaca87-2bff-41fd-8f7e-860ed42a450d · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:34.859075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:34.859075Z digest=sha256:472255d75833effd41cbb52eaf4d8325b4543f55e4f08c2d3ad6672cc5d7dffd

Observation 93bc56a0-3331-482d-9605-ce882d4f0f45 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, re- view, and open problems,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning A survey on offline reinforcement learning: Taxonomy, re- view, and open problems,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.305797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.870208Z digest=sha256:a800c24f09e3d2ebaefb0feb5a930bef761444dafb321876850511b9beb8f2c4

Observation 604314b4-3e27-4d90-8ca3-6f19d7d0d8a0 · outbound

This paper cites Beyond uniform sampling: Offline rein- forcement learning with imbalanced datasets,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Beyond uniform sampling: Offline rein- forcement learning with imbalanced datasets,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.292915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.877293Z digest=sha256:dccdc0d96ed14ca6c77577ec3177109653dd499042bb41561c09217062c827e6

Observation 236357e8-6ef3-4a9a-977a-73bbac6958dd · outbound

This paper cites The Importance of Pessimism in Fixed-Dataset Policy Optimization.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning The Importance of Pessimism in Fixed-Dataset Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:34.883014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:34.883014Z digest=sha256:e17d5641c678c6ff741e46b0a3ee4a53d9d7ce0792c65425826fc9a96e221549

Observation a8e25237-783c-40f6-ba56-904d2be06f7b · outbound

This paper cites A minimalist approach to offline reinforcement learning,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning A minimalist approach to offline reinforcement learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.280383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.888164Z digest=sha256:112cf73f9623fd02a0a3392399d264af28cd96e4884dcfd19d7aca7ada7354ea

Observation e73a3f8a-57a8-4dd5-9eab-d581031347f3 · outbound

This paper cites Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:34.892886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:34.892886Z digest=sha256:3f4598c0c5128af91c468eb5cc859bcee9dc420998e93cbd4556dfea0910cd5d

Observation b35471d9-ad49-46ca-9c13-3370f8286c4e · outbound

This paper cites Spectral normalization for lipschitz-constrained poli- cies on learning humanoid locomotion,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Spectral normalization for lipschitz-constrained poli- cies on learning humanoid locomotion,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:05:35.119772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.898184Z digest=sha256:228841a16bda318f6fafb3a3b81d91075c77445f4e50aebcff776dc13b5a203a

Observation eec35a7d-9d51-4cb5-bf0b-4ba930e96ecd · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Off-policy deep reinforcement learning without exploration,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.265944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.902501Z digest=sha256:ce99d4128aacad2b4b2fa51d6e9ad783806a19ded53509aec3abf75d059c471d

Observation c7862269-d9a3-4488-95fd-24c8c5f651ca · outbound

This paper cites Conser- vative q-learning for offline reinforcement learning,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Conser- vative q-learning for offline reinforcement learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.252104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.906893Z digest=sha256:4cfe85cbd2d4158225ab16935ea08106f4e2aa4c08d1f930ebbd4c9c46b24add

Observation f5c72a33-f7f0-4801-b359-a790c5ca075e · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified Q-ensemble,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Uncertainty-based offline reinforcement learning with diversified Q-ensemble,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.235358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.911028Z digest=sha256:ba68e003131ac99ef6a00ae97158078ad59fc0ed0f57c00ddea3a6e710c56fc4

Observation 204d28c3-6802-4fa4-b287-10233ce27d3a · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:34.915030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:34.915030Z digest=sha256:579ce04166af6d171cd355ecff50b3f91fac5769425d2f928aafca406bce2561

Observation 5be6ea27-521a-474d-bce6-e720a79165cd · outbound

This paper cites L2c2: Locally lipschitz continuous con- straint towards stable and smooth reinforcement learn- ing,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning L2c2: Locally lipschitz continuous con- straint towards stable and smooth reinforcement learn- ing,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.220917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.919350Z digest=sha256:df2396de41e756f0b0b3dcdccae4f0c84b63db4e6423d92a8efc6571499afa9f

Observation 965f766b-652e-41a8-aa8a-b01fac3cd49c · outbound

This paper cites Monotonic value function factorisation for deep multi-agent reinforcement learn- ing,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Monotonic value function factorisation for deep multi-agent reinforcement learn- ing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.207639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.923283Z digest=sha256:8c993ffd6072f0bf8043908c9e46bc1bf1f9d2192360008b64492c1704f645ca

Observation 9028cea8-d25d-45e6-aefa-462185e76c73 · outbound

This paper cites Optnet: Differentiable opti- mization as a layer in neural networks,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Optnet: Differentiable opti- mization as a layer in neural networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.193515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.927396Z digest=sha256:a05778d0381b0da3d019238d9e3df197159952bacf1d3a980b48f2582d1de965

Observation c596da04-9ee7-4e72-9769-113a3ebad4c7 · outbound

This paper cites A theory of regularized Markov decision processes,.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning A theory of regularized Markov decision processes,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:05:35.179110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:05:34.931288Z digest=sha256:a2fb48cbfb918327c22187b336058b325b53a6e39a204f28099f882f1c96284f

Observation 29ba232e-c4c5-4f17-979a-8bad596bdc84 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:34.935087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:34.935087Z digest=sha256:2c0c1bdd2b554bccc5ad9e2a8b48aa7112358ed12ebb80b6bc30a020ed9417f8

Pith citing papers

Observation afe72333-69ab-4fb1-9db2-b6f349c12b10 · inbound

Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets cites this paper.

Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.650231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T20:52:08.062450Z digest=sha256:60ffee6815a2181bc99f6d88d032537855de3c2bf70187815225c7f01bcaa794