Pith. sign in

Paper Citation Record · LEDGER

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.22008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22008 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:18:42.180146Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d7dec4f-6150-4ed1-994b-60799a054b25 · outbound

This paper cites It is an open-world city simulation originally proposed by Sestini et al.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning It is an open-world city simulation originally proposed by Sestini et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.932432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.702652Z digest=sha256:c817919cd9cd89543a51c9cc4d13fd29c4eac00d70800e01ede42275e58e53c1

Observation 8fd18dfc-2dfc-47e6-bd2e-4db9c7b22d91 · outbound

This paper cites Semi-supervised reward learning for offline reinforcement learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Semi-supervised reward learning for offline reinforcement learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.627088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:40.147376Z digest=sha256:e6532c8394d6e6c4ebb311e4340b2ddb5aac8feb1204a469c6f7e67943968228

Observation 712750b4-bcd7-4a39-903d-6e8105fc559e · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Imitating Human Behaviour with Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.482865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.482865Z digest=sha256:44757a1f9701cd580d65c17b88d141d913a052c41d7705ce39de1e4648eb266c

Observation db243e78-9b73-4234-934d-fe74ed65f3cc · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.659299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.659299Z digest=sha256:8111eceb2b9b2fb630e68ce3daa5a1ce5e09c839c51e37228e76483887683a63

Observation 1b91c065-8c05-4f75-97eb-340212b64dd2 · outbound

This paper cites Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:18:42.411597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.072225Z digest=sha256:01a54073833fb53d5593a7bcab3e0b1757de85102b54034be63989bd9f8da5a6

Observation bfbfab27-e9ab-46be-8d03-b18b17121187 · outbound

This paper cites Offline Learning from Demonstrations and Unlabeled Experience.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Learning from Demonstrations and Unlabeled Experience

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:41.211864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:41.211864Z digest=sha256:bb692d89169c30230b9503982db4fb9ab79ba89ffa89910c53bec574b22131f1

Observation d0fc60be-e3a7-473b-a2d5-a12719025707 · outbound

This paper cites Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.829576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.318621Z digest=sha256:2b76360d328be2c0e616928191c2e7ed52e87f4cc2087b815eed991d4c27de8c

Observation bf7933e9-9537-4676-b198-4e8b16331d02 · outbound

This paper cites As we describe in Section 3, it is especially suitable for offline settings.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning As we describe in Section 3, it is especially suitable for offline settings

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.521428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.444412Z digest=sha256:31ec2bbdd7a76923b8ef760a71abc013750bf952a5efbea4d835fbfcd352dffa

Observation 42e8ad3a-8bb7-4dbd-a2bf-00e9e752eedc · outbound

This paper cites All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.250384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.581591Z digest=sha256:d96ed6372844bf4baf1c70c816898da211d4af79cbbd9cb2d7e8b6cbc87a5374

Observation 23f83ca3-7541-496e-b363-2e8a34dc3744 · outbound

This paper cites An episode is marked a success if the agent reaches the goal before the timeout.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning An episode is marked a success if the agent reaches the goal before the timeout

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.646139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:41.855809Z digest=sha256:e26a6ec7a6e971897456177b70fedea5c0d1873600f8e04b411ec87959c288d4

Observation 59acd6bf-03da-4eaa-96a4-29a66e2f455c · outbound

This paper cites We provide more details about the baselines in Section 4.2 of the main paper.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning We provide more details about the baselines in Section 4.2 of the main paper

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.400993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:42.053775Z digest=sha256:9e82077e0cb9a96e9401b12e85362842d9b620f1a9c0d3ba5c6c1c7b0d24cc68

Observation f204ddd1-e672-44b6-95c5-8bafe6d88323 · outbound

This paper cites an unresolved cited work.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:18:43.133188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:42.180146Z digest=sha256:ca5560c261544275a9430b0cca90fe51464392aba13bc8b5bf5e0db3fc23f19e

Observation d4c60b3a-dd2e-4220-9dd3-04619ddab4d2 · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Efficient Active Imitation Learning with Random Network Distillation

Reference 1995

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.862234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:39.654517Z digest=sha256:f77e1b1b263877c91d598a5ce8007d44fc22887caf06b7f5e3c640476753bab3

Observation 0468167d-bdc4-4e3a-8b6a-eb989e5aefb8 · outbound

This paper cites The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:39.898864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:39.898864Z digest=sha256:39c2b8fdb061706932ceabd0be4338710172c4a501293e1204b14224b2c11d18

Observation 31bad089-20d2-4f3c-b523-8c9ce563d6b3 · outbound

This paper cites Technical challenges of deploying reinforcement learning agents for game testing in aaa games.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Technical challenges of deploying reinforcement learning agents for game testing in aaa games

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.697343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:39.752822Z digest=sha256:7c4ad6677477821137ed6cbc5a703701c5e5ddb6a371fe50e5e2aa34be6de311

Observation d0f97daf-17e3-4e84-9b5d-5f5a97969c03 · outbound

This paper cites Demonstration-efficient inverse reinforcement learning in procedurally generated environments.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Demonstration-efficient inverse reinforcement learning in procedurally generated environments

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.468020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:40.791397Z digest=sha256:d6d262f449f2d1b359803d8d8c05aaa11a011c0e85d39f94b4777ec0b896e9c3

Observation 213c6f6d-44c1-4820-bc60-a440fd8fe231 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.309764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.309764Z digest=sha256:09110804f6f96c5b3cf0492c89d179437f61f6685980acfe0fb870665225934b

Observation 2209042a-0467-4f47-87c4-06c0de714503 · outbound

This paper cites Towards informed design and validation assistance in computer games using imitation learn- ing.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Towards informed design and validation assistance in computer games using imitation learn- ing

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.099069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:18:40.929974Z digest=sha256:82bd61a22a6433550c61e3d664e35597e7ee5af60585bc2a17e2caf60beed5cc

Observation 3ebf8c51-d315-40b4-bddf-d33506d93911 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.007912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.007912Z digest=sha256:e5be5a0ac521001513fd18188fa0e12330098793e40ed64fb6f8f29f3c19ae01

Pith citing papers

No inbound Pith citation observations are available.