Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T19:23:10.328625Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.04412.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T19:23:10.328625Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 86ca782d-3e33-40d3-9bff-62a237458a54 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Online difficulty filtering for reasoning oriented reinforcement learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d87c94e-8486-482f-a55d-a04a067b6410 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Curriculum learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c3a10c-c837-4e45-ab52-cdb093e1daac · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f201116-9926-4dba-8371-ccc8f0416cd6 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7525426-0c1d-4d56-9d41-89c4128777ad · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe5363c-8cd6-43b6-8511-d624128d2c81 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a067c88-8e7d-4ff9-b524-31d4db693b3e · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Advancedif: Rubric-based bench- marking and reinforcement learning for advancing llm instruction following.arXiv preprint arXiv:2511.10507, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec597410-04d9-4162-9c9e-1b79e06bf902 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Distilling the Knowledge in a Neural Network
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b3b8f8-8bce-4f17-8ae9-08b994026120 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Large language models are reasoning teachers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10135b71-3735-4e28-ad15-dad414f60591 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71a1ee3f-722b-433a-b011-e4cf3f80c624 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Vcrl: Variance-based curriculum reinforcement learning for large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2bc3ba-3117-405a-b674-af0cc9a9f8b8 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Followbench: A multi-level fine-grained constraints following benchmark for large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f84151a-9e60-47a2-96a0-3e0a6fc49322 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Prometheus: Inducing fine-grained evaluation capability in language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0fc07a-24b2-428d-b102-4531fd40fac9 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Language self-play for data-free training.arXiv preprint arXiv:2509.07414, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0e0874-f0df-400b-a20f-d5bfa6ccf4f9 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Gonzalez, Hao Zhang, and Ion Stoica
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e86ef38-5da0-4936-ae8d-2890fc3262d7 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL SPICE: Self-play in corpus environments improves reasoning.arXiv preprint arXiv:2510.24684, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7f806a-675f-4eca-a3ba-3f1f26204c5d · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Openrubrics: Towards scalable synthetic rubric generation for reward modeling and llm alignment.arXiv preprint arXiv:2510.07743, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89443e7-8cb3-4988-a80b-b78964317481 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL G- Eval: NLG evaluation using GPT-4 with better human alignment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5687cb72-c2ef-4600-b05e-73c575c5c4d6 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Aligning with human judgement: The role of pairwise preference in large language model ev aluators
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55804b4-9c82-41dc-bb7c-b7f481b1ef26 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983e9400-539e-405b-866e-a1491cdbc756 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Teaching small language models to reason
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1278c5ee-1690-430c-bdbc-3e3d2874e800 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35c9b4c-b09f-4ce6-8843-6f0de04c0504 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Infobench: Evaluating instruction following ability in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bce996b-8755-4747-b4dc-2fe50057cfac · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Direct preference optimization: Your language model is secretly a reward model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa8ab28-6e57-4d51-8fd6-89f7137d66ab · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f65330f-e8bd-4e57-8198-5616dd2a66d6 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1053f9bb-bfb5-4c7a-aa61-bd5c0dfc8d9d · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9b0ce8-4367-44b0-8725-92dfd438c0cc · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Hybridflow: A flexible and efficient rlhf framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e24633-ec37-4cc6-b08f-13e96808deae · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Learning to summarize with human feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c52ea9e-8729-4ada-85cb-f5d7487a4626 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Sutton and Andrew G
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b52b05f8-23fd-4f93-9e5b-f95633968524 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Proximal curriculum for reinforcement learning agents.Transactions on Machine Learning Research, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb5304f-16d3-423e-bbc4-f6616d4033a6 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Checklists are better than reward models for aligning language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be063aa9-ba5f-49f8-9873-c5a00b58fc70 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Williams
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337e28fb-23ad-4217-b57b-4dd1cc1dcc5f · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9f4e00-9704-4ff9-85f3-e8357cc351ab · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Qwen3 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff97028-1fd7-4a5f-a61d-4d5f6b572086 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004c5abb-9941-4284-85b1-43c7f583578b · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88526c1a-bb01-437b-b255-0f591bb5aaa7 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Wildchat: 1M chatgpt interaction logs in the wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252353ac-9d46-432b-a971-9ed7b3db867e · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Xing, Hao Zhang, Joseph E
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d79dff-b85a-4d60-ae93-6e56fe66c968 · outbound
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL , and the criterion might be
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.