Pith. sign in

Paper Citation Record · LEDGER

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.22008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22008 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:18:42.180146Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d7dec4f-6150-4ed1-994b-60799a054b25 · outbound

This paper cites It is an open-world city simulation originally proposed by Sestini et al.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning It is an open-world city simulation originally proposed by Sestini et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.932432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.702652Z digest=sha256:ada862d689665b781f73190b9dffa3c671d14fd82425ccb6062f38fcc6a5798a

Observation 8fd18dfc-2dfc-47e6-bd2e-4db9c7b22d91 · outbound

This paper cites Semi-supervised reward learning for offline reinforcement learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Semi-supervised reward learning for offline reinforcement learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.627088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:40.147376Z digest=sha256:469ecee3f6d087d21daef458c74d0c16d0165eb1dcdf15b0044fb618a9fe0abe

Observation 712750b4-bcd7-4a39-903d-6e8105fc559e · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Imitating Human Behaviour with Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.482865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.482865Z digest=sha256:44757a1f9701cd580d65c17b88d141d913a052c41d7705ce39de1e4648eb266c

Observation db243e78-9b73-4234-934d-fe74ed65f3cc · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.659299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.659299Z digest=sha256:8111eceb2b9b2fb630e68ce3daa5a1ce5e09c839c51e37228e76483887683a63

Observation 1b91c065-8c05-4f75-97eb-340212b64dd2 · outbound

This paper cites Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:18:42.411597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.072225Z digest=sha256:d49936a34f772ebacaf4c70a4aaa8f62d1c7dfc8df003a2020ebd170f9dad87b

Observation bfbfab27-e9ab-46be-8d03-b18b17121187 · outbound

This paper cites Offline Learning from Demonstrations and Unlabeled Experience.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Learning from Demonstrations and Unlabeled Experience

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:41.211864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:41.211864Z digest=sha256:bb692d89169c30230b9503982db4fb9ab79ba89ffa89910c53bec574b22131f1

Observation d0fc60be-e3a7-473b-a2d5-a12719025707 · outbound

This paper cites Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.829576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.318621Z digest=sha256:f12557ccf0433f65772a53c77bb6afddc3e5a4f69aa06a0745c2e4143233643e

Observation bf7933e9-9537-4676-b198-4e8b16331d02 · outbound

This paper cites As we describe in Section 3, it is especially suitable for offline settings.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning As we describe in Section 3, it is especially suitable for offline settings

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.521428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.444412Z digest=sha256:b27cc883ac94bb2c4266ae7c5f83e44539a7134512347cb3d8d334ad2b83a92e

Observation 42e8ad3a-8bb7-4dbd-a2bf-00e9e752eedc · outbound

This paper cites All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.250384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.581591Z digest=sha256:30b68e077393a739c09a56f75d6ecef5af7b3df842ba194aba040c4e3be20c7d

Observation 23f83ca3-7541-496e-b363-2e8a34dc3744 · outbound

This paper cites An episode is marked a success if the agent reaches the goal before the timeout.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning An episode is marked a success if the agent reaches the goal before the timeout

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.646139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:41.855809Z digest=sha256:f2e1f45433f9038f2ca00a8d1c6d6ceafcecb027b65296fe91d17aa10babb17a

Observation 59acd6bf-03da-4eaa-96a4-29a66e2f455c · outbound

This paper cites We provide more details about the baselines in Section 4.2 of the main paper.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning We provide more details about the baselines in Section 4.2 of the main paper

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.400993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:42.053775Z digest=sha256:ca027076f9d6d872bd3722716302ec03abff0c2f30649451042febb43e1dd1e1

Observation f204ddd1-e672-44b6-95c5-8bafe6d88323 · outbound

This paper cites an unresolved cited work.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:18:43.133188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:42.180146Z digest=sha256:2dbc9621966eb502b2d67ac630c0f449ada79db60b018d205caca80bfe624f07

Observation d4c60b3a-dd2e-4220-9dd3-04619ddab4d2 · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Efficient Active Imitation Learning with Random Network Distillation

Reference 1995

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.862234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:39.654517Z digest=sha256:b5a04f1985cd923dc5dad5d972a775d785ab3eb2e5a560d53e6b8f3f7dae4f72

Observation 0468167d-bdc4-4e3a-8b6a-eb989e5aefb8 · outbound

This paper cites The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:39.898864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:39.898864Z digest=sha256:39c2b8fdb061706932ceabd0be4338710172c4a501293e1204b14224b2c11d18

Observation 31bad089-20d2-4f3c-b523-8c9ce563d6b3 · outbound

This paper cites Technical challenges of deploying reinforcement learning agents for game testing in aaa games.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Technical challenges of deploying reinforcement learning agents for game testing in aaa games

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.697343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:39.752822Z digest=sha256:43650e5f329b5e0c8424235d55476fee3ea9b47361c38ba96dd0f6e6cfbcf015

Observation d0f97daf-17e3-4e84-9b5d-5f5a97969c03 · outbound

This paper cites Demonstration-efficient inverse reinforcement learning in procedurally generated environments.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Demonstration-efficient inverse reinforcement learning in procedurally generated environments

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.468020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:40.791397Z digest=sha256:1b93c43ff6bc3a3612a62eabe4b73f1a07093420ee5be3f0e843cf7a0c3644ed

Observation 213c6f6d-44c1-4820-bc60-a440fd8fe231 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.309764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.309764Z digest=sha256:09110804f6f96c5b3cf0492c89d179437f61f6685980acfe0fb870665225934b

Observation 2209042a-0467-4f47-87c4-06c0de714503 · outbound

This paper cites Towards informed design and validation assistance in computer games using imitation learn- ing.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Towards informed design and validation assistance in computer games using imitation learn- ing

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.099069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:18:40.929974Z digest=sha256:da2f336c3cf995d64ff9103820a318699d4a3dc1ac03d95b93774d469dba1d9c

Observation 3ebf8c51-d315-40b4-bddf-d33506d93911 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.007912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.007912Z digest=sha256:e5be5a0ac521001513fd18188fa0e12330098793e40ed64fb6f8f29f3c19ae01

Pith citing papers

No inbound Pith citation observations are available.