Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2411.05193.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:57.529911Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.183441Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a5e731ea-e33b-4d1a-b2ac-aa81f4b0a581 · inbound
ShiQ: Bringing back Bellman to LLMs Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97dd5a8-7a18-4ff5-b2f4-fbb033ad99b5 · inbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d4ac9f-b08a-458d-abdd-506c69e3831d · inbound
Post-Completion Learning for Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c91798-d270-49d5-8e48-98e58360054b · inbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bfc2373-26c3-421d-bc93-94b66b9ae816 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.