Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2303.00957.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:40:10.017433Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:06.217026Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2809a598-64be-4929-b590-d4fc542cf60b · inbound
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d193b40d-bb8a-4dc6-90b7-c540a4f1cc73 · inbound
SimulPL: Aligning Human Preferences in Simultaneous Machine Translation Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · inbound
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400ac5c5-b05d-4e79-8f8c-9218e0d3ccca · inbound
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a91461f9-1259-4997-b09b-baf9131b7999 · inbound
MAPL: Multi-Objective Preference Learning for Robot Locomotion Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 15fc3cd5-f288-4405-864a-c698ba0c5770 · inbound
SPLC: Social Preference Learning for Crowd Robot Navigation Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c5e12a0-2ffe-4969-81fd-e524c9241314 · inbound
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 227
Source-reported events for the cited work
Unavailable: canonical work link unavailable.