Pith. sign in

Paper Citation Record · LEDGER

Magistral

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 44 inbound Pith citation observations for arXiv:2506.10910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10910 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.848892Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:38.122268Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 58df91f7-f6d0-43f2-914c-f803515f43e7 · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

Magistral What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:22.151310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:22.151310Z digest=sha256:51f6098dba6649fdd3ace8f6e0cf5bfc7673df46aaacdd6bebd8aafb3845a172

Observation 5c547a00-d3e3-4bef-9df6-0750e017b308 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Magistral DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:22.289997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:22.289997Z digest=sha256:6bebee06b0959590714c08be4db9a9c917c4d83b71665847069b981a3af95286

Observation 5704b324-819e-4fe2-bbb4-d084dc1041c6 · outbound

This paper cites Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures.

Magistral Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:29.137514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:22.472941Z digest=sha256:62501d901cfc0f191c9148b4da8db94c1bb25566e0f8a4198d71fbdfa5db54fc

Observation 67e25e8a-2b82-49b5-921b-5f515a8f23fb · outbound

This paper cites Polyglot Benchmark.

Magistral Polyglot Benchmark

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:28.977000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:22.624332Z digest=sha256:b7fb386fa12383c3615c3f3eb3565ee277b2a216fd1a9e74760d3df5bb42221a

Observation 5003029f-9c06-4414-9506-4461b56d06f7 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Magistral OpenThoughts: Data Recipes for Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:22.778940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:22.778940Z digest=sha256:6933ef8c995f39eedf39c13caf65aa9827c32a973832874f3e40ab4a7eee13e2

Observation 68a4518f-83be-443b-9663-c4830f70049d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Magistral Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:22.981703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:22.981703Z digest=sha256:9e1800cefed97a1c577a484cb9ab2430e1845255a5e8d54c38c391557c34b2d2

Observation 8d724e63-8a2b-4c53-9039-7a7c68ed16e7 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Magistral OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.130127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.130127Z digest=sha256:beffe935afab0e94391b0416a8fb2e1269ca8f41665fef3b9cbc8ca2e45e40aa

Observation 3aaefb17-6335-4a84-9917-cbe598a1b973 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Magistral Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.278363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.278363Z digest=sha256:314e25a83cc2e4d94a6ff901bedb52f31a336003f5c02a5f37fb14afa368e872

Observation cbf1b255-9e5c-4119-80eb-1f58bd63ee53 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Magistral Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.479130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.479130Z digest=sha256:36ac24ba2b98892cd664bb53b3ad4f33030a0106ab5b5c3774e236b42befa0a2

Observation 48ba2a9f-8a07-4070-8b93-7762621a965a · outbound

This paper cites OpenAI o1 System Card.

Magistral OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.630431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.630431Z digest=sha256:255096e3dc1fbcfecda91087fd2e97984b2f9bedf6b8ab077bf3b07e1db82e86

Observation 0b17f5c1-93e3-476c-9a23-20a2d86b9f3f · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Magistral LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.789584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.789584Z digest=sha256:c8506cd29b4e71138c8fa20a8a4782de868febd0a6c089abadea17404ff3f374

Observation f778baa5-f855-40d6-b5b6-d2f99d9e1d71 · outbound

This paper cites FastText.zip: Compressing text classification models.

Magistral FastText.zip: Compressing text classification models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:23.974974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:23.974974Z digest=sha256:e6f8a905ae282332a8783d8331604ba6310a57fe64e3dfafca6ffc2be14f13e6

Observation fec18fcf-db8d-4007-8074-085886f07043 · outbound

This paper cites Visualizing the Loss Landscape of Neural Nets.

Magistral Visualizing the Loss Landscape of Neural Nets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:24.080434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:24.080434Z digest=sha256:f88b60a1fda49e7f107f5d96650ad521979c853b0d733867e52ebce6a79eff49

Observation b9445add-fb95-4f97-8b48-afd2089ea90f · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Magistral Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:24.249365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:24.249365Z digest=sha256:e913ff92a65d430a124de9bd0bc74a95ddbf9b6fe4421def74f71b233a989a1e

Observation 628e55f3-111d-414e-b23d-9926e437097c · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Magistral Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:24.433838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:24.433838Z digest=sha256:268f7ceac10c4a5d5eaf8faa45bbde030e683b561f50fdfa0ff635794fa6720d

Observation b9cbcabf-32ce-4767-bc51-ebc7e5781d15 · outbound

This paper cites Mistral large 2.

Magistral Mistral large 2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:28.649078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:24.559640Z digest=sha256:102887bd113d35ed1981f48097cf59f2c3bf84a62fda0ce2bf78b40fcb71d0a3

Observation ba51863e-03b3-4038-9708-26edd0658d47 · outbound

This paper cites Mistral medium 3.

Magistral Mistral medium 3

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:28.343947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:24.717927Z digest=sha256:5971f2072c952bf1dbe41bb4249f9b3471a90aa5365132ddb6ec9474e86b9f89

Observation 80499156-04c0-4aec-b60a-9c656f3b4258 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Magistral Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:24.817910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:24.817910Z digest=sha256:29ec5e25a2d9faa6f14aba0aa1b3f8a96975b3a70427a7965b5a9996c349d430

Observation 24591ab7-adab-4bca-adea-f95306197332 · outbound

This paper cites Codeforces cots.

Magistral Codeforces cots

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:28.009074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:24.973926Z digest=sha256:cb975fa43409f6c3e1b3fb2e43960682dfc0bfef19c0e8ef7923449399953b54

Observation 4c805a10-7964-4e0b-9367-b218e3a0d53c · outbound

This paper cites Humanity's Last Exam.

Magistral Humanity's Last Exam

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:25.101520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:25.101520Z digest=sha256:60622b62ffb2db754ff44acb7318699fc2eee431e9e2699f15b6213ea1d53b0a

Observation 21d0f88c-bac0-4a20-ba78-ad10edf48079 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Magistral Gpqa: A graduate-level google-proof q&a benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:25.209936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:25.209936Z digest=sha256:117a9ffbda3866621c508f37752bef97e5ab9dbe154cc79ee28d8dc006d88110

Observation 948f4ce8-b893-48e3-a349-082b1c6adf57 · outbound

This paper cites Iterative methods for sparse linear systems.

Magistral Iterative methods for sparse linear systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:27.763019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:25.393789Z digest=sha256:62e98aaf6fe593fd91d30864a876e4c7efe9fb0adcc8eb4f69f84c2530987b20

Observation d77bcf82-acb8-45a3-8aff-7a87174bf9c7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Magistral Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:25.552219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:25.552219Z digest=sha256:58399f5933f6f6c1c10ad2e1fe07200f7196e501d9b6dea1ed59274ca81f6d97

Observation 6a9d1a23-b828-4223-af6c-82725fa9ac0b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Magistral DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:25.739088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:25.739088Z digest=sha256:b56136519f4e079de760a4aaab200350471bdbaa014b27954848bb879cdf09f1

Observation 7f4314a9-ef00-41f8-afec-0260a50552f4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Magistral HybridFlow: A Flexible and Efficient RLHF Framework

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:25.902064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:25.902064Z digest=sha256:a28f03fc7920112471eaa576804c253c95b4d46d5870942cd7092767fda3f394

Observation b0e1489a-96a4-46bd-90f8-18784a3d8458 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Magistral Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.065118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.065118Z digest=sha256:33ede240d664f6f0071f318d885e586c663f6ae001d195e83319f8fce14814cb

Observation 62aa2849-4bc7-4bce-953f-b36661f1c2ce · outbound

This paper cites LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training.

Magistral LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.241394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.241394Z digest=sha256:6ebd9b8f505d691c4f77ba349d8b6ae7bd4eb7bc3d2e50f621e1b68e689125f6

Observation df998082-8d39-4586-bd26-e31f6501b6fd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Magistral DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.415218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.415218Z digest=sha256:1c051d09a4ff0fa2c6c2627fe2bfb54937c2766def76ce748a729ac460d28b74

Observation afb421ca-e8f2-4607-8461-078a57556093 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Magistral Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.563646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.563646Z digest=sha256:49364f104e3741c469afcfd2186e5956520deff2c4cd243417afe973c393b506

Observation 7f2c72e9-0ff2-4342-a537-1c5365d8db3d · outbound

This paper cites Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark.

Magistral Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:27.582378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:26.723792Z digest=sha256:a91c29c1eb119f9116a6a3032cfe0d31e69f84cb96a761b4fedc34ed851aa97c

Observation 52cca6b6-2d53-4db9-9b01-09a147b99b98 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Magistral Instruction-Following Evaluation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.848892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.848892Z digest=sha256:c80c802cc6d76ee6bd383b9ae2fe43c4119aa161efa92b0fb6bb0b54f22181f2

Pith citing papers

Observation 55adf272-4e07-4ef0-975f-904dc0245558 · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.560228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.560228Z digest=sha256:c7f4b7336d32717679c5df4db158109a83a42dbd24e660cface7ec8705f192f2

Observation 47fa10e6-114b-4b08-b746-7b13c998cd2c · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Magistral

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.564942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.564942Z digest=sha256:8dcca1066882600c2ff204fc2766fbf4996989a042e926b21ccb62c827d70e4e

Observation b779c2bb-1bff-4a9e-b9c9-72297ed03027 · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients Magistral

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.588485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.588485Z digest=sha256:5bde8b6180328a23c0533ae8fa626adb3630de4e7fecb23fbe371eb06b72f481

Observation 7edf36cd-5d18-4ee0-97ca-7ebc8605bc18 · inbound

Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency cites this paper.

Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency Magistral

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T01:02:46.136737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:02:46.136737Z digest=sha256:d0947c36ba4e1f5e73c8b7fb60b36abf0be58bdb2ed9dc32520d1723a1015ee3

Observation 45628e2c-bc46-4f8f-9750-044521932656 · inbound

Hermes 4 Technical Report cites this paper.

Hermes 4 Technical Report Magistral

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:54.579019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:32:54.579019Z digest=sha256:29bf918ff887d198b06e783a51e6ffc3b96bff370b0170c31c91baaeb6db6a48

Observation a6555b84-43a2-475f-b906-cd989fed6db3 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report Magistral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.459947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.459947Z digest=sha256:73b8db8b2220ead9af2d018588ef685f264cbda53f8c2fbaa243e5494188c700

Observation ba52a9bb-f820-45e3-a236-89c70f7539ee · inbound

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing cites this paper.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Magistral

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.122268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.122268Z digest=sha256:d6eba957db54a106142340bffa81bcd67e0060fa51f07ffbd53660804fa1adde

Observation cd458219-c9f1-439e-ba58-d2f6ef8fa086 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Magistral

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:38.099803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:38.099803Z digest=sha256:6719b057397667467854ab5489ce89639fcc6e735c2cb21fa5b014eb8f068c3f

Observation eb683d59-e733-45c2-b54a-fd10b31088b4 · inbound

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems cites this paper.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Magistral

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.580795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.580795Z digest=sha256:9927ba718b090c47cad6c81feadfd46af6250cc517a4967134392e2f40736252

Observation 10e694eb-7163-4a4d-a0cc-93b0f75ea00f · inbound

Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements cites this paper.

Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements Magistral

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:22.409827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:47:22.409827Z digest=sha256:d239aa62f35d7a95ad6c466f6152a930ae7451092f678384c926ded37ab06a78

Observation 4a56cd11-dfa7-40d7-833f-20687921b302 · inbound

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation cites this paper.

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation Magistral

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:40.861861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:40.861861Z digest=sha256:46bdfd8f3143a835269bf933baf8d29104bde27020e9235591593e49b1a826a3

Observation 6bfc45b6-5981-4322-8ba9-320755b5e74a · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts Magistral

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T20:54:21.599958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:9692b30f9dde07688533a22c1e85fe7242a1a43c23d4ed01f5050752bbc5ef98

Observation f6298168-ee72-4d00-936c-212358e8e52e · inbound

Are Large Reasoning Models Interruptible? cites this paper.

Are Large Reasoning Models Interruptible? Magistral

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:07:52.637167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:07:52.637167Z digest=sha256:f4cc58e53894a74a1c869572253b366b744affad7745ace5ba5e04d5b67f899b

Observation 8f14ebbe-2e79-4187-a375-14693ba94948 · inbound

Ministral 3 cites this paper.

Ministral 3 Magistral

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.787396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:2c8f6be9d69085e6e56501bc0f313d95b03dd6566e44907d2578805f22901bc7

Observation 44d95e98-01b7-446c-8e6f-a5982282210a · inbound

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models cites this paper.

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models Magistral

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:54:31.109524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:54:30.992080Z digest=sha256:81271062cc90a02bc40a5999a9c7ef0e173eb48dd71681c87592983d57cb4e2f

Observation 66e62282-7e92-4464-b7fd-5efd6a4a2270 · inbound

Simultaneous Speech-to-Speech Translation Without Aligned Data cites this paper.

Simultaneous Speech-to-Speech Translation Without Aligned Data Magistral

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:57:01.017891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:57:01.017891Z digest=sha256:a55dd5c172950dddffe67e81e1c4c9e70044db4b243de09699c47d14382d13b3

Observation 7f4e80d6-e4f4-42ed-ba6b-3d5c09ab1a63 · inbound

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling cites this paper.

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling Magistral

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T23:14:39.407159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:14:39.407159Z digest=sha256:8759d154c40f1b96248d59e8322ae166b17b165607d08578826b451e9d1d1034

Observation 2f7f2d9e-b2fd-4310-9f85-63e6be00e0ed · inbound

Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis cites this paper.

Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis Magistral

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T20:07:28.015727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:07:28.015727Z digest=sha256:052fc71a0691a1fdafe5894d01d53ec5d0a492ac0d10bd9b0d37647700266504

Observation 4c23408e-dc3c-4095-a984-bb2e6ebd51b7 · inbound

A Novel Hierarchical Multi-Agent System for Payments Using LLMs cites this paper.

A Novel Hierarchical Multi-Agent System for Payments Using LLMs Magistral

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:40.456510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:40.456510Z digest=sha256:d917e5828da9e7e613beae0b0c87977430b92b4da46cd8f44c7bab3b3d370ac2

Observation 44806c9e-e9ac-4a7b-916d-e5f39cb499c4 · inbound

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics cites this paper.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Magistral

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.999962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.999962Z digest=sha256:b22562c3221cbf756a077333d824c2b60cce65004bef495ee7fe1b813ae87d09

Observation fee22f91-491a-47a8-8cf6-b9927dcdba70 · inbound

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios cites this paper.

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Magistral

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:14.494409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:05:14.494409Z digest=sha256:704bed387b6e5c0981a33bf3b1a7ee0d4633c6fc70964995bce3b47dd9a26736

Observation 8a26308b-c52a-449b-b4aa-de02d9f608a0 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Magistral

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:29.394776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:7f9f1b447e30bbbe9b8dc6f67cd803c192f68448c41ad938112abf80311f426f

Observation 54215b74-247f-4322-b231-219345a6c660 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards Magistral

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:27.020410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:cf5a577f4efcef304febb5147cd27f041b1addcf2be561b06b75222431aa4a97

Observation 034b9bc8-cef4-4c22-ae27-69e9705552b9 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Magistral

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.473510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:995c2ab1be2edec295c544ecdcc466af90245983d48ab098a1596a78ec3bdb83

Observation fe227479-5036-4c4e-9c4a-efa2eac81788 · inbound

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards cites this paper.

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards Magistral

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:45:21.193991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T04:40:52.854907Z digest=sha256:c012d6799b402100ff5ae0be671bc96659384b3b3f6d7d9dd1293a4410ce4612

Observation 58c87324-1eda-4b9c-9218-514685fecf62 · inbound

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search cites this paper.

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search Magistral

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.595526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:19:18.121220Z digest=sha256:b7b7f4a2593db4c2d6cdb0b5cc5df2c17e4463b7e05b1e7381bcdf1e730a72bf

Observation 68576317-78a5-4df6-8851-12825ea438f8 · inbound

Language as a Latent Variable for Reasoning Optimization cites this paper.

Language as a Latent Variable for Reasoning Optimization Magistral

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:31:06.613349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T21:40:37.499246Z digest=sha256:1cd5847885e279d0a89c9eef4379303a0e20b2b74d00278d44faed5660cff396

Observation 163e5329-0c8b-416f-b23e-05fd51ef83ca · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Magistral

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:15.966281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:ef4405822caa563dd952ae37427eb09a37af91dae39e86e5382b37585664883a

Observation 3288c138-cde4-4090-805b-18756e46e300 · inbound

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models cites this paper.

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:06:06.496149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T18:51:59.402906Z digest=sha256:e566453800fa068528156939e1295e9c73363ed6a5cecf77c076bec9fe05e3bd

Observation 79c6f8be-3c78-4651-b24a-d1ed5273043c · inbound

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models cites this paper.

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:55:31.085847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T07:48:28.381924Z digest=sha256:2f8542ed4f35af08b4cbaa3a19135781d90b06e69929a72d892889135374c4f2

Observation a26de0a7-cdf0-4b87-9f9f-cda7263cef3d · inbound

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models cites this paper.

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Magistral

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:33.269530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:20:33.269530Z digest=sha256:f5e195d340fca79a21cce5f9ea82eff5ad1dc8558822ab2d0ad26784f03389a0

Observation 8ba3c1fa-e843-4348-b259-15b6efe9335f · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models Magistral

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:43:21.816674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:413493e0ffbb207bc865b5f917d66d664fd93d4b8f4c3969362a6137ac28578c

Observation 3b6793d2-7bf7-4271-80e8-b6900abc0ea9 · inbound

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning cites this paper.

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Magistral

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-05T10:20:57.186398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-05T10:18:18.717871Z digest=sha256:865868ba1835fb03bdd3ed067889fdaec16bde5dd9bee212d30b3200ceec0857

Observation 1b858ce7-4e44-4d77-a08b-8ca5bbb14aa8 · inbound

A Primer in Post-Training Reasoning Data: What We Know About How It Works cites this paper.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Magistral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.508069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:413591e15f2a3268c3b82573250ff9a202dcf21fc34d94094c8a4ec5acabb20b

Observation a2212d16-e9e4-4e84-9124-6dc20dae2279 · inbound

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs cites this paper.

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs Magistral

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:13.147843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T22:08:45.988451Z digest=sha256:4473baf7cdb371e6777f0d3f4ede293d7039c160fe2113b93f42db4eb620fbdf

Observation 25d944ad-0225-4817-aa40-72d6a6956ee1 · inbound

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval cites this paper.

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval Magistral

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.158738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T03:35:30.841853Z digest=sha256:39f9b6367685c969c5e18060feec940dec276de3351f2fd80f04e9bf3a1a08c3

Observation 3658b2d5-21c7-4ed9-b73f-9483efdbd014 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Magistral

Reference 133

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.307984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:7cf39c6cf8af5c1c71022b7e069c0a7080be35963b8232f967e523823cd8b2ad

Observation 91a0ab2b-7c07-45aa-9722-061ead9fff09 · inbound

Predictable GRPO: A Closed-Form Model of Training Dynamics cites this paper.

Predictable GRPO: A Closed-Form Model of Training Dynamics Magistral

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.603748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T06:40:47.911021Z digest=sha256:283af033b3f169f23078e4d9a5c9a11c7a1b8be0b414cca8c74a5db567006585

Observation c4f07051-995b-4380-ae9c-81404c659281 · inbound

Predictable GRPO: A Closed-Form Model of Training Dynamics cites this paper.

Predictable GRPO: A Closed-Form Model of Training Dynamics Magistral

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:21.687480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T20:26:15.667646Z digest=sha256:4472314aa1102a262307d86003dc23aa35a26c11c4220f3f10e54139f08b7197

Observation 7dbec6d2-9bca-4ca5-9c09-a0e8a27e7e55 · inbound

Cost of Reasoning in non-English Languages: A Case Study on Japanese cites this paper.

Cost of Reasoning in non-English Languages: A Case Study on Japanese Magistral

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T14:12:44.286699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:12:44.286699Z digest=sha256:ea8f31bf10ca4fc2f762b92b41d911e953397b8d323bfa28b10161c57a915d15

Observation 1941e2df-6267-42b0-98bb-c2f57dc50f90 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness Magistral

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-01T08:36:30.534949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:36:30.534949Z digest=sha256:203aeffe81039b8045ad908b00ab4100184bfd50e98cdbf21765ba0686ee4f77

Observation b9076e27-ef59-4b11-8172-d2e50da442da · inbound

Contrastive ESA: Human Evaluation of Multiple Translations at Once cites this paper.

Contrastive ESA: Human Evaluation of Multiple Translations at Once Magistral

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:50.895042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:48:50.895042Z digest=sha256:01dcef4abde8a07287c870ad0c0cc96c1ca4a717cd87932f8a2ce1594af256fb

Observation ebb49bd1-9cb4-4867-8d6d-5f47806d1d42 · inbound

Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability cites this paper.

Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability Magistral

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-04T10:56:20.082102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:56:20.082102Z digest=sha256:a86a053af14f7093eb224a3d990a4e495e3b88e8c3aeb1fabe06aac2eb0a030c

Observation 63f9a6ff-2722-4a98-bdb1-3722d656b668 · inbound

Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning cites this paper.

Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning Magistral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:11:01.995426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:11:01.995426Z digest=sha256:3ca1cff2c0abc9484de2de6d344557b365856f9c906d14963389543c736cf2b8