Pith. sign in

Paper Citation Record · LEDGER

Offline Safe Reinforcement Learning Using Trajectory Classification

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2412.15429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15429 v5

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:30:06.718481Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d392448-e575-462f-a099-a19e14447871 · outbound

This paper cites Additionally, SafetyGymnasium includes five velocity- constrained tasks for the agents, Ant, HalfCheetah, Hopper, Walker2d, and Swimmer.

Offline Safe Reinforcement Learning Using Trajectory Classification Additionally, SafetyGymnasium includes five velocity- constrained tasks for the agents, Ant, HalfCheetah, Hopper, Walker2d, and Swimmer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:06.938861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T11:30:06.718481Z digest=sha256:6a023279f968d308adcec424c24e1fe4c311b002ddf491d37ac2b49e00505f47

Observation 18c774c9-cf56-4bd1-9c01-f50a1989c6ae · outbound

This paper cites Information asymmetry in KL-regularized RL.

Offline Safe Reinforcement Learning Using Trajectory Classification Information asymmetry in KL-regularized RL

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:30:06.866006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T11:30:06.675090Z digest=sha256:046501268f2913dc9d5e5f189ccaf38a73f2ea716b8c43a797c646307e012619

Observation 68046617-dae9-4b99-92c5-85eef071965c · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

Offline Safe Reinforcement Learning Using Trajectory Classification Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.691008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.691008Z digest=sha256:87daf1dd1b003ddca68479f7140b39bac3d65e9f2e25aeaf5837c89d20374128

Observation 4a019047-0e98-42fb-8f3e-472e92eea9e5 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Offline Safe Reinforcement Learning Using Trajectory Classification Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.696490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.696490Z digest=sha256:b2b16cbf6924e01064cd2089462808a5b3b6ccaf611f8718ab0ee2e3de939d71

Observation a1dbd71a-4b92-46be-99c3-d0db40b902b8 · outbound

This paper cites Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model.

Offline Safe Reinforcement Learning Using Trajectory Classification Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.707943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.707943Z digest=sha256:6398688647278857eaea3b454f1ebe5acd0c4c53c12c4eb09da815a001e63a4f

Observation 2fe4b687-9f5c-4d0f-a8b4-67613f411bd7 · outbound

This paper cites In Aaai, vol- ume 8, 1433–1438.

Offline Safe Reinforcement Learning Using Trajectory Classification In Aaai, vol- ume 8, 1433–1438

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:06.956003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T11:30:06.713348Z digest=sha256:cc35cce2462f5c503c13fc79ea5d8912343e3d67e5555c89a639e9e684923921

Observation ca64a2f1-7689-4d97-94bb-72e9c128350c · outbound

This paper cites Reward Constrained Policy Optimization.

Offline Safe Reinforcement Learning Using Trajectory Classification Reward Constrained Policy Optimization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.701825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.701825Z digest=sha256:24428e3c344615a814ba39e6941f65debea7514c379344f9dd8418f9b3a9b122

Observation b42f9f0c-697c-4533-95c8-63d9743e628e · outbound

This paper cites In ICML, 2052–2062.

Offline Safe Reinforcement Learning Using Trajectory Classification In ICML, 2052–2062

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:06.974156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T11:30:06.670083Z digest=sha256:c581715c6458f7ffa5d9b80b3d28495d3ead7cb618c6268b8bd1c937833bce93

Observation 0a912bba-8465-4c31-ae85-d0b710dc8def · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Offline Safe Reinforcement Learning Using Trajectory Classification D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.664950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.664950Z digest=sha256:ca63f5814f0eed858001ef804aeebd69bc5ec299fd0a46b047711b75adfebc23

Observation 4e11c989-1b4f-494d-b00d-2a18c3091576 · outbound

This paper cites A Primal-Dual Approach to Constrained Markov Decision Processes.

Offline Safe Reinforcement Learning Using Trajectory Classification A Primal-Dual Approach to Constrained Markov Decision Processes

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.654064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.654064Z digest=sha256:722ba3b1c27efb86d46af885b21f4c2ebe32a34465ae79d60308833fc33a1760

Observation 8e6c6fc6-97fe-4f05-9690-46f1434fec3a · outbound

This paper cites A Review of Safe Reinforcement Learning: Methods, Theory and Applications.

Offline Safe Reinforcement Learning Using Trajectory Classification A Review of Safe Reinforcement Learning: Methods, Theory and Applications

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.680248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.680248Z digest=sha256:10bfe529d81633f61201c7f22e521bf057cd633f4d6bebd131c8543172e5c82e

Observation 9ee70c93-8be7-428c-8750-68b977ed3737 · outbound

This paper cites OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research.

Offline Safe Reinforcement Learning Using Trajectory Classification OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.685475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.685475Z digest=sha256:25af2976654470ab1199264a1a17cf8972e65f08815e88ea723d179b39dd7d4c

Observation a060bc4d-1403-4f04-a38f-6a69ccc63fe3 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Offline Safe Reinforcement Learning Using Trajectory Classification KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.659787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.659787Z digest=sha256:92b41ed54219142e27281ac9c4f108991b897c3827d076bc8ec41b7152b29832

Pith citing papers

No inbound Pith citation observations are available.