Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:52:02.517746Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.25659.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:52:02.517746Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d4e5bf94-41ca-4fb9-ac42-97162986f8ed · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in neural information processing systems , volume=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea19546-555e-4372-899e-faad9d654d27 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in neural information processing systems , volume=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031f23cc-6ff5-4363-88f5-4a3d83f983e0 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5f8084-4cd4-4beb-9bb4-a4af4091f56a · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029dbbdc-9bc9-4769-91fe-1a7632f31b9e · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume =
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14cc2846-ea2e-46d3-9b93-9cfa3adf2951 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c13eec-aa26-4575-87fb-3084278f5d59 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406497d2-7e78-4100-b276-e90df5358498 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aceaeceb-60e9-490c-b384-bcbd9465245a · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34dae7ad-51d8-4702-aaec-601b67a86ddf · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization NeurIPS 2025 Workshop on Efficient Reasoning , year=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63151e3e-2ef8-4355-b3a7-d6476766c662 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c41472d-e77c-4984-9a62-0c9dc467881a · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f59b9b5-f567-4b8a-9de0-ec738d5f571e · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 156cd443-06da-4e8b-9036-9fdb0d85d21f · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4c47de-57bd-4b61-95e2-ec02196a1e89 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques , pages =
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d706cb-a763-44b6-9d35-af692cb66590 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Machine Learning , volume =
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d368c68d-316b-41d8-b882-bed880420339 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190e7da1-0325-4fa0-b071-885335bda75a · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Findings of the Association for Computational Linguistics: ACL 2024 , pages=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3ea0cf-2511-4549-8d46-98bbe0eec267 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2065375-3491-475d-9581-1a15bbd250c4 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization From Generation to Judgment: Opportunities and Challenges of
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b870122e-ed2f-431a-ad5d-a06161f4a113 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fcc690d-0329-4642-a660-b6d0647ddd8b · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Computational Linguistics , volume=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8f0b2b-8204-4f70-9228-e9241234c450 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Constitutional AI: Harmlessness from AI Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f63742b-9734-4084-9bc8-1ea3a7a3d81c · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c04bb7-a48e-4f65-b89a-eb4165b2b3d0 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Probabilistic Attribution For Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b284e0-dce9-4b4a-ac38-72c9538955f3 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2512.23457 , year=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b0cc37-69f0-4e28-9e7c-54d4f13ed042 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2511.10507 , year=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df0d396-1817-46a0-bea1-0533f77635f5 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2509.19199 , year=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f318d49-7162-4062-9395-5e70f20c5c66 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72dedb9-1533-41b7-a921-bf81ed9c770e · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2603.08754 , year=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647ab12c-787a-448b-a255-14fc95088eab · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2601.08430 , year=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a5e376-83be-4e00-9059-fc6f8fb46862 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104e30d9-1c2e-4632-abe1-f7ec2a78079c · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2508.16949 , year=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d82b2ed-2b5b-4b33-8bfc-2f0cbbea53c8 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Self-Distilled RLVR
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f85128b-f8cb-4594-8b9c-a692bf7224ca · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6523b596-bfc5-436d-8191-23449ca69c77 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61fdaf94-133b-48ec-91c5-1134e1a81d18 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Instruction-Following Evaluation for Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f26d11b-1a5c-4d90-842e-fa713e36e00b · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed24d3b-e25f-4c7f-860c-7829083f532e · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Group Sequence Policy Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395710e0-aa83-40c0-ab23-0e61267842fc · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization The Llama 3 Herd of Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ba7d64-5a45-4870-97dc-af7f0c83f7ad · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Qwen3 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5bfd7a-f17f-48ca-88a5-d674ff3233ad · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Qwen2.5 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b357547-316a-48f5-a73c-4ec5b56ddc10 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc21fb3-ccbe-4119-b851-cc5e0a8819cf · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfeeab2b-5a52-4fe3-847c-679898dae623 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8610db-4088-4402-8f66-f47e301b87b4 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2510.00194 , year=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0c09e4-a157-4e5d-8442-b0d670056806 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization OpenAI o1 System Card
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1d62ca-9229-4710-aed5-547a1c4786c4 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Nature , volume=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8027817f-d979-4dfd-9b7e-7779049f0268 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb902165-c989-4696-a123-a16c86dcfbe4 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the Thirty-Eighth
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292ad71d-5ead-4706-aeb1-4d505edcaef5 · outbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the Thirty-Eighth
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.