Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.848892Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 40 inbound Pith citation observations for arXiv:2506.10910.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.848892Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:11:01.995426Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 58df91f7-f6d0-43f2-914c-f803515f43e7 · outbound
Magistral What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c547a00-d3e3-4bef-9df6-0750e017b308 · outbound
Magistral DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5704b324-819e-4fe2-bbb4-d084dc1041c6 · outbound
Magistral Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 67e25e8a-2b82-49b5-921b-5f515a8f23fb · outbound
Magistral Polyglot Benchmark
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5003029f-9c06-4414-9506-4461b56d06f7 · outbound
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a4518f-83be-443b-9663-c4830f70049d · outbound
Magistral Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d724e63-8a2b-4c53-9039-7a7c68ed16e7 · outbound
Magistral OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aaefb17-6335-4a84-9917-cbe598a1b973 · outbound
Magistral Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf1b255-9e5c-4119-80eb-1f58bd63ee53 · outbound
Magistral Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ba2a9f-8a07-4070-8b93-7762621a965a · outbound
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b17f5c1-93e3-476c-9a23-20a2d86b9f3f · outbound
Magistral LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f778baa5-f855-40d6-b5b6-d2f99d9e1d71 · outbound
Magistral FastText.zip: Compressing text classification models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec18fcf-db8d-4007-8074-085886f07043 · outbound
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9445add-fb95-4f97-8b48-afd2089ea90f · outbound
Magistral Understanding R1-Zero-Like Training: A Critical Perspective
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628e55f3-111d-414e-b23d-9926e437097c · outbound
Magistral Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9cbcabf-32ce-4767-bc51-ebc7e5781d15 · outbound
Magistral Mistral large 2
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba51863e-03b3-4038-9708-26edd0658d47 · outbound
Magistral Mistral medium 3
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80499156-04c0-4aec-b60a-9c656f3b4258 · outbound
Magistral Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24591ab7-adab-4bca-adea-f95306197332 · outbound
Magistral Codeforces cots
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c805a10-7964-4e0b-9367-b218e3a0d53c · outbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d0f88c-bac0-4a20-ba78-ad10edf48079 · outbound
Magistral Gpqa: A graduate-level google-proof q&a benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948f4ce8-b893-48e3-a349-082b1c6adf57 · outbound
Magistral Iterative methods for sparse linear systems
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d77bcf82-acb8-45a3-8aff-7a87174bf9c7 · outbound
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9d1a23-b828-4223-af6c-82725fa9ac0b · outbound
Magistral DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4314a9-ef00-41f8-afec-0260a50552f4 · outbound
Magistral HybridFlow: A Flexible and Efficient RLHF Framework
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e1489a-96a4-46bd-90f8-18784a3d8458 · outbound
Magistral Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62aa2849-4bc7-4bce-953f-b36661f1c2ce · outbound
Magistral LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df998082-8d39-4586-bd26-e31f6501b6fd · outbound
Magistral DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb421ca-e8f2-4607-8461-078a57556093 · outbound
Magistral Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2c72e9-0ff2-4342-a537-1c5365d8db3d · outbound
Magistral Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52cca6b6-2d53-4db9-9b01-09a147b99b98 · outbound
Magistral Instruction-Following Evaluation for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55adf272-4e07-4ef0-975f-904dc0245558 · inbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47fa10e6-114b-4b08-b746-7b13c998cd2c · inbound
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Magistral
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b779c2bb-1bff-4a9e-b9c9-72297ed03027 · inbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edf36cd-5d18-4ee0-97ca-7ebc8605bc18 · inbound
Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency Magistral
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45628e2c-bc46-4f8f-9750-044521932656 · inbound
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6555b84-43a2-475f-b906-cd989fed6db3 · inbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd458219-c9f1-439e-ba58-d2f6ef8fa086 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Magistral
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfc45b6-5981-4322-8ba9-320755b5e74a · inbound
Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts Magistral
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6298168-ee72-4d00-936c-212358e8e52e · inbound
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f14ebbe-2e79-4187-a375-14693ba94948 · inbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 44d95e98-01b7-446c-8e6f-a5982282210a · inbound
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models Magistral
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66e62282-7e92-4464-b7fd-5efd6a4a2270 · inbound
Simultaneous Speech-to-Speech Translation Without Aligned Data Magistral
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4e80d6-e4f4-42ed-ba6b-3d5c09ab1a63 · inbound
The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling Magistral
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7f2d9e-b2fd-4310-9f85-63e6be00e0ed · inbound
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis Magistral
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c23408e-dc3c-4095-a984-bb2e6ebd51b7 · inbound
A Novel Hierarchical Multi-Agent System for Payments Using LLMs Magistral
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44806c9e-e9ac-4a7b-916d-e5f39cb499c4 · inbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Magistral
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee22f91-491a-47a8-8cf6-b9927dcdba70 · inbound
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Magistral
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a26308b-c52a-449b-b4aa-de02d9f608a0 · inbound
Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Magistral
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 54215b74-247f-4322-b231-219345a6c660 · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards Magistral
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 034b9bc8-cef4-4c22-ae27-69e9705552b9 · inbound
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe227479-5036-4c4e-9c4a-efa2eac81788 · inbound
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards Magistral
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58c87324-1eda-4b9c-9218-514685fecf62 · inbound
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search Magistral
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 68576317-78a5-4df6-8851-12825ea438f8 · inbound
Language as a Latent Variable for Reasoning Optimization Magistral
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 163e5329-0c8b-416f-b23e-05fd51ef83ca · inbound
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Magistral
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3288c138-cde4-4090-805b-18756e46e300 · inbound
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79c6f8be-3c78-4651-b24a-d1ed5273043c · inbound
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a26de0a7-cdf0-4b87-9f9f-cda7263cef3d · inbound
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba3c1fa-e843-4348-b259-15b6efe9335f · inbound
Self-Supervised On-Policy Distillation for Reasoning Language Models Magistral
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3b6793d2-7bf7-4271-80e8-b6900abc0ea9 · inbound
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Magistral
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b858ce7-4e44-4d77-a08b-8ca5bbb14aa8 · inbound
A Primer in Post-Training Reasoning Data: What We Know About How It Works Magistral
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a2212d16-e9e4-4e84-9124-6dc20dae2279 · inbound
The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs Magistral
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 25d944ad-0225-4817-aa40-72d6a6956ee1 · inbound
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval Magistral
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3658b2d5-21c7-4ed9-b73f-9483efdbd014 · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Magistral
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91a0ab2b-7c07-45aa-9722-061ead9fff09 · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics Magistral
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4f07051-995b-4380-ae9c-81404c659281 · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics Magistral
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7dbec6d2-9bca-4ca5-9c09-a0e8a27e7e55 · inbound
Cost of Reasoning in non-English Languages: A Case Study on Japanese Magistral
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1941e2df-6267-42b0-98bb-c2f57dc50f90 · inbound
Training Large Language Models for Self-Explanation Faithfulness Magistral
Reference 145
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9076e27-ef59-4b11-8172-d2e50da442da · inbound
Contrastive ESA: Human Evaluation of Multiple Translations at Once Magistral
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb49bd1-9cb4-4867-8d6d-5f47806d1d42 · inbound
Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability Magistral
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f9a6ff-2722-4a98-bdb1-3722d656b668 · inbound
Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning Magistral
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.