Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:06:49.755721Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2508.06026.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:06:49.755721Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T14:00:28.509134Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6e2097c0-2f18-4025-83f8-884c8e571991 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Y.; Rajbhandari, S.; Awan, A
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3acfe968-e085-46da-a75d-1714196f7471 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future PaLM 2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fee6bf4-7908-49a4-9541-4fcfcb5529af · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbfa7b49-95de-4eed-9f27-e478b0537b14 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71f2658-3540-4f16-be15-29d0564b1ef0 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3546670-28f0-428a-a813-b2636864ef32 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future PaLM: Scaling Language Modeling with Pathways
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17403456-8f5d-45bb-b58a-84c21c0c5e2d · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 999e88b1-31ff-4cfb-a966-595b267ba0bc · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future D.; and Chatzoglou, P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c99e57b6-dc00-4d09-b4a1-c0c4d0d6ae38 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Improvement in Language Models: The Sharpening Mechanism
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a32b81c5-329a-439f-92a4-83d5b77c13e4 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Large Language Models Can Self-Improve
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f747ef8-1b49-4793-a342-84c2e7d649cb · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ba34a6-ad6a-4d61-9c6c-372f29c13ac3 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66377ed1-f64a-4ec4-8053-682d072536f6 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future o pf, A.; Kilcher, Y.; Von R \
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 292e2e37-2247-4dfa-a57e-561bb553f050 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future H.; Gonzalez, J
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b83f706-b805-4cad-9357-506fc9a9a37f · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Diverse Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12c86bd-478d-4814-947b-c2d2814803e3 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Generative Judge for Evaluating Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e30402c-ed2e-40f9-9e2c-01d45a429b1c · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b48448-a52a-4f71-b12e-4937390791ac · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Alignment with Instruction Backtranslation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef6c809-ae6b-413c-b9af-9d701bc46b00 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27ce3375-f590-421a-ae23-8a0be3fc2df3 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac02ce60-fad8-4221-a813-a20db1a9a42d · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 092382b6-2b8c-42d5-9650-e3796259d9ac · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d08cd5a-6a5d-4a26-b6d3-1e817c31283b · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27153f03-a1cb-4a4b-aeea-4333db13b026 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4b7da3-019c-4df1-8e25-37eb52b1eecd · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c7bb6c53-4d80-4d3f-9489-cb20d2990fa9 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future D.; Ermon, S.; and Finn, C
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972df6f0-fd7c-42db-bcdb-8dd0c67eac67 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future D.; and Arora, S
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef979ca-8cf6-4927-9d87-3c4045b782f1 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd4722a-e41b-4475-af7e-ca396a61768c · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future LLaMA: Open and Efficient Foundation Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0106e4-d14d-464d-8503-afc5bcdd59eb · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 19a68efa-94f5-40ea-a1fb-beae266ab74a · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future CREAM: Consistency Regularized Self-Rewarding Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60045f1-012f-4a39-93aa-e7081a442b2c · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Enabling Language Models to Implicitly Learn Self-Improvement
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71b2bf67-f471-482f-ba00-5427082fa2ee · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbbf387-75a1-4980-98ba-75c7278a89f1 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Qwen2.5 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f206878-43d7-44dd-b0a8-1b6e50f83983 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Teaching Language Models to Self-Improve through Interactive Demonstrations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 596f078d-45fb-46bf-a7ad-bcc89012defb · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Rewarding Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3111235b-0f77-4cf0-8a98-0f23c4a51f66 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future GLM-130B: An Open Bilingual Pre-trained Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92bd15e-a55f-47c6-9899-d9fa0b7ae95a · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cefe01cc-2bf6-412e-9050-1c1a2d005fcc · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Process-based Self-Rewarding Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15701eec-a68a-4602-bb80-7ceeb2824b31 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Language Models are Universal Embedders
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db368bf3-22c8-401f-b634-de41099e2368 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Forcing Diffuse Distributions out of Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee15eaf-15d6-45bf-9c0d-29fae194ec06 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18db4afd-80b6-4da9-b0f4-c8e5b09750d9 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb64a4ed-7654-425f-94c7-47c635f5722a · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future , " * write output.state after.block = add.period write newline
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f03bad6-0778-437a-8257-e1538fe59399 · outbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future write newline
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be78f0cd-54ad-4ce9-9b6a-dfeb43182ab9 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.