Pith. sign in

Paper Citation Record · LEDGER

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2412.08880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08880 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:33:37.721942Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 484da714-e99a-4520-b896-f63ea4633c16 · outbound

This paper cites Constrained policy optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.528192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.528192Z digest=sha256:10ba9ab8eb8ba33f96f796e22957e366be42f16e1df4052df928e63b89826bd9

Observation 614de274-5e77-4761-80bf-c28644564dc3 · outbound

This paper cites Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.307332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.534971Z digest=sha256:364dc7e01db0a5b7045901776bee32b68402665421b411d8c2b209b95bfd6d39

Observation a791a083-db0d-40e6-93c3-9d19b8d87889 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.289432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.540639Z digest=sha256:0d26734fd26b9fecbef04abd5f9db320cc3950a259a6949e2de08d789acf8489

Observation 1db5df65-eda1-4c92-9873-86c4032d64d4 · outbound

This paper cites Constrained Markov decision processes.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained Markov decision processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.547006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.547006Z digest=sha256:8ffb752c9aa39e906ef0cbee0b566beb3ee7c9b391414ddea76e1f1f6d12234f

Observation 636cafb2-0f7b-4233-ad75-f1032fce1913 · outbound

This paper cites Convex optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.552930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.552930Z digest=sha256:58f356ff959a65afccf780be4ea9d75460815a6e930a8e7acceed3aefef5e7a8

Observation 93c8ef20-0422-4cfb-86d2-28ba9d322c2a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.558683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.558683Z digest=sha256:ddcaf4f52059adc81a94b8d656003177d66062968af24e616c17a2a960d4bae6

Observation 6b258255-db0c-43f1-a7bd-bd18cb70c4f6 · outbound

This paper cites Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.564869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.564869Z digest=sha256:fb6b96b85eca058983b8aa5ad673dced9bfd764a7288f671465e71899669a2b2

Observation 9256cbd1-13c6-4efa-8f2c-b3a2e828c268 · outbound

This paper cites Directional differentiability of optimal solutions under slater's condition.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Directional differentiability of optimal solutions under slater's condition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.226032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.569920Z digest=sha256:2b334c8cd79646d63461df4dc1996d7d9a3a76cecb1133cc92290ff72d2e0e98

Observation 71a64338-d0d0-4560-ad92-c457ed052236 · outbound

This paper cites Parenting: Safe Reinforcement Learning from Human Input.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Parenting: Safe Reinforcement Learning from Human Input

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:38.005117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.574512Z digest=sha256:d79781725dcd5675eb2b42b5631e83f6482398c03088f81792563beea9e8ddfc

Observation 0deffe96-4c7c-4b64-bbfd-34b362135362 · outbound

This paper cites A comprehensive survey on safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A comprehensive survey on safe reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.580345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.580345Z digest=sha256:64faae394f06019f4e83cdcd9292a68e74883b5cc00756766c99068b51777720

Observation 1d1f6f5a-aa67-4549-8be7-782a782f91a4 · outbound

This paper cites A primal-dual augmented lagrangian.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A primal-dual augmented lagrangian

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.199580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.585226Z digest=sha256:9d021836eb7d8018a8b4327091270cc876f3af79b231f844890f2a5c8fb1e15d

Observation c38cf8b3-58f2-4f41-9428-ffadf8c3f167 · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.592440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.592440Z digest=sha256:4131fd4d3e419cf184c4fc6f59efcb9c716f68f54e4f6840353d996582a94fb3

Observation 8643a124-d913-4c15-9b1b-750249d9db3b · outbound

This paper cites A Review of Safe Reinforcement Learning: Methods, Theory and Applications.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A Review of Safe Reinforcement Learning: Methods, Theory and Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.597399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.597399Z digest=sha256:acdaea41c3938e04c5da21c5c6063e65bdaabbe3cf59c0ddd6eb9752a40304ce

Observation 5ee64d99-f8cb-4f9c-9fd2-da6482429ed3 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.602358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.602358Z digest=sha256:d00c11703660ce10667c8479f2dd153cd3661af31bafef2795874ace54c72eb3

Observation 866f1926-a3d6-47a9-aa85-735a1e3990d0 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.606897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.606897Z digest=sha256:6c1e884c9cc5771f594a68e40ba41a67ce1a765e9ae03b4d9a8bd1bf8b19cd21

Observation a0e80708-02c8-4cf3-8516-cc6f0c593526 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.612195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.612195Z digest=sha256:664d20aa0551fd82cfa917af330392c8dcbf57c58ae5da9ea8110d05d036ac9f

Observation 6f5cde42-7188-4a64-b4b7-cf4fa14bf058 · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.152117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.617099Z digest=sha256:dd83be8962c279c52cbfe205843421e2df9765aef6efc27709eff3ee3f87dabc

Observation 14972fe8-b442-4a3d-9485-393ac1488e1b · outbound

This paper cites COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.622251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.622251Z digest=sha256:dca0286c0f5560d675340198186de2e1949bf8f6e82ada9bdaa281d89dd588e2

Observation f31ca234-8c65-48e2-ae24-d88c464fbdfd · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.628179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.628179Z digest=sha256:4b68a91425f1973f694e0d01a124f043933918dab2efef7456ef01085a2b24fa

Observation 33bf6911-2cbc-4a33-befc-e5bf0747f0d2 · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.634783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.634783Z digest=sha256:f1ef72b26d095a015c769049af855a80ef42b50d43c33dc29c9e9f32ffae5aeb

Observation 1f0345e2-16b3-4acb-9857-da80b41b1e4f · outbound

This paper cites Constrained decision transformer for offline safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained decision transformer for offline safe reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.137434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.640165Z digest=sha256:f8bddfda79415cd79e8bcf795830feae880f082fe91081c51659f174cc9800b2

Observation 16d8aabc-e554-40fc-a42a-dc56489c2632 · outbound

This paper cites Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.645535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.645535Z digest=sha256:b3add74a79a2f36022f0c93653587c18db9c4323b9984a089b16fcddc9f3cba3

Observation cffaf56a-76eb-496b-89f8-c54cafa1ff6c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.650823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.650823Z digest=sha256:3846afcf44cd6fd3897f53fa0fa34e7574131acae1dca9e1ca94dbf6fef70a2a

Observation a21cab1a-23d3-4346-9d9e-152408a0540a · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.656127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.656127Z digest=sha256:516a224dc24e3b15736e60aabed27dea2fb93eb7e0c3dab50cad7d7ec9bfcfeb

Observation e4adc729-f884-4777-bc6c-968f6d854617 · outbound

This paper cites Trust Region Policy Optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Trust Region Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.661644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.661644Z digest=sha256:ac3fae083e58bdf7b2b7ddb944464bd3910ed3f217c6e3e4ad390ab8edaab4d0

Observation cccddb50-a740-4495-9b42-15adee25b7c5 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.667394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.667394Z digest=sha256:3d658f4263c4a967d76200c01371300e2c12771451e062e5cfcb36bfe4edbc0a

Observation 3e162c79-8e30-482d-a879-d855d690bc2a · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Responsive safety in reinforcement learning by pid lagrangian methods

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.120946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.672685Z digest=sha256:6158913c6c7b557cfaeeca8b4feb8c58e726a95fa372c76165439fd78da348a3

Observation 6216db71-c306-4fbf-8ee6-4d198cb6b999 · outbound

This paper cites Constraints penalized q-learning for safe offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constraints penalized q-learning for safe offline reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.101582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.678677Z digest=sha256:e0821345b16919acbe7139f4443c1a994d38bfd08e0c7d4e488a43f1b6474ac9

Observation d0a69486-bf64-44aa-aa04-b430122cc777 · outbound

This paper cites Primal-dual stochastic gradient method for convex programs with many functional constraints.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Primal-dual stochastic gradient method for convex programs with many functional constraints

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.085356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.684038Z digest=sha256:f270524b9edc7efef60b3a119bcbd1ea41d6f8f142136aa902db7393f8f0e27d

Observation 9eabaaed-dbf7-44b9-86f4-8d62f22dc0ca · outbound

This paper cites OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:37.808510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.689199Z digest=sha256:b4416da1022cabdb59fc9967a450fa5dc41b66e32d094128b804acde30d56780

Observation 0e665c63-23ba-468e-a65e-3daf84174d1d · outbound

This paper cites Reachability constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Reachability constrained reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.067425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.695599Z digest=sha256:d1330348d4343e1225c4dbbadbebe8bf20dd8eff9a30c9bd60ebafdf905381d0

Observation 8f5e93cf-c4d0-48de-afad-450c830f1720 · outbound

This paper cites Penalized Proximal Policy Optimization for Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Penalized Proximal Policy Optimization for Safe Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.700628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.700628Z digest=sha256:00193541696d80d87d91b92f6d13bbfb88c48dce650816b9e6a89f67824a5abc

Observation b5486b46-e347-47f7-9837-461c07eaa0ff · outbound

This paper cites Evaluating model-free reinforcement learning toward safety-critical tasks.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.050910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.706021Z digest=sha256:a39e58554b4a6bf4af0390eceed47b1bb5c8692703a71a3d9c894f10c621f66e

Observation da699856-f6a8-4d7e-b7eb-eca65762b5af · outbound

This paper cites First order constrained optimization in policy space.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning First order constrained optimization in policy space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.034039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.711703Z digest=sha256:5885536b9ffe50adc72ace86f7b8137b7ea03a84df12e536085113f86f7442b0

Observation b59b261b-3528-4132-90d9-c51b8f8f0880 · outbound

This paper cites Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.716533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.716533Z digest=sha256:8a25e3f25aad1f0ba68d6ed462723b09f73589f339d91e87959db956d6a1ea51

Observation 0dc13ffc-13f9-44e5-94dc-f1d0ab8f545c · outbound

This paper cites Maximum entropy inverse reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.721942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.721942Z digest=sha256:487ae24e9ef2927d07877e0709b02b87cd6d722127e011b34f61a6b157464f6f

Pith citing papers

No inbound Pith citation observations are available.