Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:52.537874Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 13 inbound Pith citation observations for arXiv:2506.14965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:52.537874Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:51:25.047158Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
58 of 58 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 3138bf82-e82c-495d-a067-60bd336cd684 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83a9bb53-ebd3-4115-8c44-f19668ce2396 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71916b87-9b50-4dd0-b789-6164b85fbc6d · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0a1db52-a0ad-4743-aede-4cd590b43bed · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fcbdfac-f309-4ad1-ab2e-818520f1635c · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Then, we need to find m + n
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 904a1ea9-7ff3-4a58-b939-a12268677ca3 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective It then iterates through each contest, checking if the rating update is allowed based on the division and the current rating
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 606bc65d-bf7f-419a-bd0f-21ad666aa33e · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8bad5977-fa9b-42e4-9682-e70077173bfa · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90db118a-a375-4049-919c-501a5cedd94a · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective 1 and Div
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e9b52ce-7861-462f-8170-1132ff8b69e2 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bd8f8df-6715-49d8-836e-75d60fa87122 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c39eab6-36d7-423b-93b8-3d96b128a135 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5a7bc77-65b7-449c-8445-e1bba119cd06 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Let’s re-evaluate the steps to ensure accuracy
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 561eaa2d-27b8-4cd3-90e8-069caecd54e0 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5b81542-0dc0-42a4-9aff-88dcc6d10f33 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8d5dc60-c242-4513-8d50-64d8f780a078 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d308de17-9e8f-45b7-826b-a31712440022 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5b18e92-3841-46a3-9fb1-638b46c2bfd6 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccf1fad5-58cd-4759-ac53-005dc8641e56 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79c8865c-3eb0-4ba2-88e1-675fa05812a8 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2df6be70-039f-445e-8351-8395f955c2b1 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c2d2e3d-62a6-40d2-ae85-dcae07540fbc · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 34521ad9-4368-4240-a31d-5f3518465ab0 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44c0796f-5a42-4105-9e5e-207ea3a16213 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8b8d9f8-62a7-46fb-b45b-320123b70775 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ceceaa9-03e0-4f6a-b932-678bdb381182 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Your task is to analyze puzzles, spot patterns, and provide direct solutions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8d9ba3b-30a0-497e-9c0c-c1e794a98bfc · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13623a4e-18eb-42b3-b338-6baf05f91d48 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dfea3190-e89e-4293-92e9-24e51bbe27de · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33c17e1c-ce9d-45fb-8a53-b9cb09664b21 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b86bc8f-0d19-4976-b28f-5a63ca37dcc6 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2431a87-37bf-480e-a23a-7bc492c49acd · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17b7b37a-ebb1-4681-987c-e4ecc4bc640e · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc407327-b6b3-475e-ba27-073ddf48cbcf · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6204e8e-b5f1-4d90-af83-7d1015add0bf · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective p u b l i c _ k e y
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31d903f8-0fa2-4c32-a3cf-612d280e13e5 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ff46a1f-bb33-4f72-ba58-af439a0d7337 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54cf4116-2865-4cde-a7f6-343a9a6b0fd7 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2dcc7f77-1bb7-43a9-871d-083c1146bc3a · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8662592c-77f6-4ca8-ac2d-de10044aa35c · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 548e4237-aed6-4703-aad2-1c1b2f0113a9 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective The prime factorization of 901 is 17 × 53, so p = 17 and q = 53 (or vice versa)
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71f0f3b6-de95-4441-9b64-4f8793f176ba · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40d3ce10-ab8c-4b2f-a6e7-ea8f0bf71a7c · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e56162b0-1cf8-47ee-be70-183e190f68b5 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e93df25-5d81-411b-99fc-5789220fec3c · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective This means 3 × 555 mod 832 = 1
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53a754e9-234e-48cd-a62f-437d649cfbef · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0cc73c5-82b0-4762-a70f-33b77bcb26fe · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea9d8924-6369-4f9a-b0ee-a2c5d2e6c470 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95f034e0-584d-42e8-9b33-bf6f4229a4ee · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective p": 17,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 011ddf69-08e5-4da5-9471-161c7b2226e8 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f50becb5-bb82-4d67-8e73-054a4081ad62 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3f10b22-49e6-4199-a844-39eeac744b0d · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 032dd5d9-aaa5-4e80-a23e-aebea949cecd · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fa4e05a-4bb8-4504-8adc-e482108b97af · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 844fa836-7563-432a-90e9-589f1cb1b3c2 · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 066e2452-4e7c-4be8-a09a-ffa35138cf4f · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84b07b2c-747b-41d3-bae5-54fa71256eed · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective More specialized applications emerged with FinQA (Chen et al., 2021b) and ConvFinQA (Chen et al., 2022b), addressing TQA within the financial domain
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c29a27ce-28de-41c8-a0e1-1ac2ecb79cfa · outbound
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27600c42-52c2-4520-8d41-a13110548b31 · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 23f41cff-96e0-403b-b367-604fc7bac8fc · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e0d3d1c-e15f-453d-833c-25d07c6ef6d7 · inbound
Dream-Coder 7B: An Open Diffusion Language Model for Code Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3e99fa-9b84-4a9a-b883-cf2046791ad6 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0336c218-8eeb-441a-85a1-478bcc30418e · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef994c77-03b6-4084-9afe-6a46a4a4c9db · inbound
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67dcaf45-0935-4531-8b0e-08546f2e69e3 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f25aea38-5b2d-48c2-9cec-93deeab9d4c9 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93fb66b6-df41-4211-955e-23dc44cf4e4e · inbound
VIMPO: Value-Implicit Policy Optimization for LLMs Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b3f1521-154f-4fb4-9be1-f1a318a894d0 · inbound
Loop the Loopies! Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76885f17-4a3b-4958-9ea8-7305911faf1c · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff86dea-41c1-47b1-9ea2-89aaca1fd3c3 · inbound
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00d57f2-0c17-4e4c-b441-2e87735c546b · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.