Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:39.945503Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.06292.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:39.945503Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e259bdb-1853-4f76-84aa-ec5ff05cb212 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59f42dc6-7d9c-4be1-adc3-97899f90b612 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f3f9575d-4960-4351-974b-6d1dd7ddc3d3 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 978bf97d-65b1-438b-94a4-ab59511d346a · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1381f10-04b2-4af2-ae1a-e2b5632dedf6 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8a67c46b-6cbb-4da8-906c-726a865b6ef7 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a44d61a-c70d-4648-b021-b38faf6ebac7 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 28b5ae49-fd87-46da-a06b-728da4083026 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b71dc55c-33ca-4308-a853-269496bca3d3 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e7c4d8-2c34-497f-b0bb-7a57312fd4e4 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7c67607-e814-42c0-a8e0-af4837770113 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d5103b4b-4463-4aac-b4bd-6b3cdf90c96a · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33448f0-f078-42fc-90be-c542b2e24317 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unsupervised Prompt Learning for Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b4b98e-2343-4dd3-9aa1-d6aa3620b9c0 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8c6a49-c328-42d9-9f32-12d33108a241 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models o pf, Yannic Kilcher, Dimitri von R \
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fd313a3d-62f6-4ce5-91a2-395096329e06 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ca02fc-c8ee-4bf9-8bcc-1864ceaa277b · outbound
Mutual-Taught for Co-adapting Policy and Reward Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9c420a-51b0-48a4-98f5-93beab23e59e · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Hashimoto
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46282fff-87fd-41ac-8404-771df9e80b75 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277eeee6-42be-4398-a7f1-b5a2eb80f71b · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d01e6eb2-b3b3-4777-bf7b-e3e8ad51cf25 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a0298e-512c-4696-87af-53bfcd238469 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c6a41b-b39b-428f-b89d-50d428430fdd · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca2ed841-9dfc-4063-acfe-bf54bb93437f · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4f1290-f09a-4fd8-bc74-bc7fb255912a · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950088f6-7ef1-41a3-b947-585f9315510d · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a06adbd8-8925-4aa7-8c30-88739ff09d19 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5becf1b0-5333-47e0-a261-560c85222f6e · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54dcd746-5496-46ad-907d-1bc06e1664b6 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c370cb21-0787-42ff-85a4-5c6ca7a02d41 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56ba907a-9370-4dd8-a591-d43550b8194e · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65e4bd1c-ee79-473c-bd2b-5549e6116acf · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570ee809-c4ba-4efc-ad84-de2f1b6863cd · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c4a7e11-e3a4-46cf-8633-c45d1a077b12 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fdaae8eb-a5b5-4938-ab43-48aa74c9bcf8 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6247927f-d8b1-4abb-a09b-43bc9f775eb8 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1935e351-7c0b-4431-828e-425bf210485d · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 20027d6c-b0f4-4dca-96a2-aa95bbc23a91 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05910a02-702f-4b74-91d3-591b3c2acc1d · outbound
Mutual-Taught for Co-adapting Policy and Reward Models online" 'onlinestring :=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96265a0b-a2e8-4105-8ca1-1e62070ec375 · outbound
Mutual-Taught for Co-adapting Policy and Reward Models write newline
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.