Pith. sign in

Paper Citation Record · LEDGER

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

As of 16 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 5 inbound Pith citation observations for arXiv:2507.10541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10541 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:31.807793Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:41:01.511851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T14:26:20.339663Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45263c56-4939-4fd5-a35f-802585d0b16f · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:30.988070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:30.988070Z digest=sha256:0ededd97e40654911dfcda4f7f043b8b5090bd88ae609fd4a4e77532be5fca3a

Observation 1a4ed58e-988f-41b8-88fb-1d88bd327e52 · outbound

This paper cites AIMO Validation AIME Dataset.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once AIMO Validation AIME Dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:34.128623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.077245Z digest=sha256:08b5104a994c2f65d38caf9a84e6c7d5bb2be0fb9ec3a540098d07c3e8440432

Observation 2746420f-7962-4a4a-8836-444cd63e6cc4 · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.201122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.201122Z digest=sha256:c583d0b6c18d706120a38abc137fe076cc4335a3432002aec5766c7cbe2729fc

Observation ebfce389-c508-42f1-9643-13c3c3f797ea · outbound

This paper cites Llama-nemotron: Efficient reasoning models, 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Llama-nemotron: Efficient reasoning models, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.325011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.325011Z digest=sha256:85e89d40787e0ff5650a20c3bbb87959cd413761ecbd3b2018e5c1ac9e71ca06

Observation 7180110f-4272-4077-be7d-091de1f211b1 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.347313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.347313Z digest=sha256:f1ea980c0b4dd56156ad4b280dec955423999710949fff075ecb96daecf95a52

Observation c161d8cb-73d7-4eb3-ae87-d10ad418edba · outbound

This paper cites Batch prompting: Efficient inference with large language model apis.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Batch prompting: Efficient inference with large language model apis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:34.104687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.355798Z digest=sha256:9614385663cd60a46190b6418b4680fb350271af480bd48fbf7def944df76e94

Observation a058e47c-cbf6-4ccb-aea2-2ad48f7d49e4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.373974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.373974Z digest=sha256:7b3ccc0989079e97f671edb185b9280715ad744df9a631bb88f7aff6833aed87

Observation 5f70281e-45cc-43cf-8302-8ba8372d1a95 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Process Reinforcement through Implicit Rewards

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.381615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.381615Z digest=sha256:d2f11fb79882b11e12b95da0208b1948c422e1a89da28ae2052410d502a51d16

Observation d41e278d-97de-4f5e-9cb1-3cbaab6e17ac · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:34.077786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.400902Z digest=sha256:ed460c335dd77ff555acd73014bee0e153a89d361276c3058c9456cd9819472a

Observation de32417e-e383-4e0a-b2b9-456d3a002b82 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.408154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.408154Z digest=sha256:dfcb0ab9f019adf65b6031f76d81de118428621a15f82f79d4524964a81c4ee2

Observation 06093aef-6291-4c7e-aa5b-3a6027b33ad1 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.417001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.417001Z digest=sha256:d497fd82283abf37cf590541e8a368be9fed9212ece4a1095adc198978ac8707

Observation bd95410d-82d2-4ac5-96a3-ae684348192f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.423282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.423282Z digest=sha256:72ebc974f29d4b395f042785985bfeea3f633ebc881d4bc52a7981ac10cabb01

Observation c2cf8936-1197-459c-ae5c-e25a64c8a1d5 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Token-Budget-Aware LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.430010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.430010Z digest=sha256:0aa336f4f9458d42c5cb82fc84595dd28f6e209ff4f77b7cf6423860837d503f

Observation 47fbd123-0066-482a-a30e-1246a88e7a83 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad- level bilingual multimodal scientific problems.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Olympiadbench: A challenging benchmark for promoting agi with olympiad- level bilingual multimodal scientific problems

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:34.047658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.435294Z digest=sha256:54c1a6cade3dae8aaee0c1ee782aab68a9f790bb80cc04b59f38d1f2c3711f27

Observation 6b20a37a-2a46-42b7-bc35-26cacc39a33a · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Measuring mathematical problem solving with the math dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:34.018469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.440805Z digest=sha256:0b5077a726b5e6e18bed1adfd7904dd879873a630fd5792d481a2de0bd582153

Observation ba63a264-19df-4168-a4f4-5a6159ecca68 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.446614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.446614Z digest=sha256:e647548c338d112788d5793589992f41973bb907913294768c5ebb18a51f7492

Observation 10fd9226-c8f3-484c-b213-a18a7980bd64 · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.arXiv preprint arXiv:2504.07086, 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.arXiv preprint arXiv:2504.07086, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.463634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.463634Z digest=sha256:4fc1d00fa2b80919c0728243e846dadd227eb3fb41a21d88f7d11e99aea40582

Observation c167a1d1-345d-4cc1-9172-f02160a6a1e8 · outbound

This paper cites Compound-qa: A benchmark for evaluating llms on compound questions.arXiv preprint arXiv:2411.10163, 2024.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Compound-qa: A benchmark for evaluating llms on compound questions.arXiv preprint arXiv:2411.10163, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.469333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.469333Z digest=sha256:fafc8c2c9e4edca0320ad01c3f92f60c01819b23149ed347aec81b97599162ee

Observation 5374261a-fc43-48f6-8a2d-89c639e966ec · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.474664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.474664Z digest=sha256:3339b2c5c2d3329b513379e5a60b41e0515466995f6f83c183372526f18b4c5d

Observation d9b75dce-a692-4a17-96c3-e4c1b56b8201 · outbound

This paper cites Qwen2.5-Coder Technical Report.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Qwen2.5-Coder Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.479733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.479733Z digest=sha256:c968f4911b86f482dbebc124de71bab881a3e6ec1391b47d3e6e01385f1a3ba7

Observation 8827ea6b-82b3-4512-8c17-6d9418c39241 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.486293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.486293Z digest=sha256:a94d97bd0a0b4d51269ac49cd3311e40ea270c10dc1a7e3a013425e1973557a5

Observation f8d00e6a-8f2d-43fe-b230-6fbfe265d4b8 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Swe-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.973252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.493366Z digest=sha256:26f0c4686b3086b9f47658dc2ec942d26c2a346d4893f79b82e0fee71acaa4d0

Observation fbf33416-9f5c-453b-89b2-20a4bfe2f7ac · outbound

This paper cites Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.498612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.498612Z digest=sha256:85582cdc559413db8a919681d197743610c0ea1eb9695a8c8c5ec8b1c9d1caec

Observation def049e9-6de0-4ab7-8a45-1410b4ad7429 · outbound

This paper cites CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.504727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.504727Z digest=sha256:ec70f766e6556322632aaffb28aef0a1885cac5010a356e37ae1b3a5388ba8fc

Observation 64f47877-0a6e-4bd5-878e-343153bdd0d2 · outbound

This paper cites MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:32.543004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.510786Z digest=sha256:e4272b51429991e3f84d35f043d637c94df3d1bf066eab8cb0859d0216db1c52

Observation 0c557471-d394-4ea3-b7cf-7aef5ebf7d00 · outbound

This paper cites Aime 2025 dataset, 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Aime 2025 dataset, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.947508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.517198Z digest=sha256:a71f19907b6bf3067e720dd8795b89ccee1378ad5d7aedffd43ef7adcec645d1

Observation b990de15-63eb-4dcd-a11e-cf6bbba9e77f · outbound

This paper cites Lost in the middle: How language models use long contexts.T ransactions of the Association for Computational Linguistics, 12, 2024.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Lost in the middle: How language models use long contexts.T ransactions of the Association for Computational Linguistics, 12, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.918264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.524111Z digest=sha256:c3dced908763cc7848f96f35b81dcbd6ea41ae5fdd3996151e9b468d76931aa1

Observation 5204b94e-7408-418d-b2d9-cda9cc9501f7 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.531602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.531602Z digest=sha256:4c51c4e3990e5bf9185e528bddfdc15d36a9e2762f01fd34f9bba8661fad2ca0

Observation 36a9691b-8521-436f-a46e-944e0bcf5b58 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.887993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.537341Z digest=sha256:bca8b5328e22545c8bce5083fccf0933ca1825434a01f780da746dfd9f0d1a21

Observation f1118e6a-569d-4c15-ab13-16284aed28a1 · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Real: Efficient rlhf training of large language models with parameter reallocation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.544110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.544110Z digest=sha256:588ff237af0b53b0abc2093af082b7cef1e9586d924daaf5069f743920fa9257

Observation 3e452e6f-618b-4eac-b43e-04b4ac0359f8 · outbound

This paper cites s1: Simple test-time scaling.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once s1: Simple test-time scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.550519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.550519Z digest=sha256:659076326069db66725919f6edd86f5cfeface982282ee384c07e1be7df7cf50

Observation d41f9fc1-155e-455a-83e2-f47493bfacc5 · outbound

This paper cites Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.557650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.557650Z digest=sha256:3481492f1eb201fcb53289b9423a38d354a98963fc610e291a1b06846240b4b6

Observation 39780489-485f-4189-bbc1-fd9062f744c1 · outbound

This paper cites Openai o3 and o4-mini system card, Apr 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Openai o3 and o4-mini system card, Apr 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.841219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.573060Z digest=sha256:545870301498b67f51a7a425796a173a58c85ddd9f0ff265f8d8f1458cabb716

Observation 0a7d6353-7a6c-4cce-9c82-e62afc5fc3ee · outbound

This paper cites LEMMA: Learning from Errors for MatheMatical Advancement in LLMs.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once LEMMA: Learning from Errors for MatheMatical Advancement in LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.578289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.578289Z digest=sha256:029cf1f9e9d6195a04eb24757fb090973bc86d31a5e67a9a53520cebdd6c70ff

Observation f3315be9-bdf9-4c7a-a7b7-0a90b1e52275 · outbound

This paper cites MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.583773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.583773Z digest=sha256:29ea94d251629b962f81208acb1d2c0eb5a5ef6c49a6b181a457b508508ce521

Observation bce54451-14c6-4b2c-aa51-696bf63c5a03 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.589921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.589921Z digest=sha256:1eedd1eb507f302e7e18dea65630484c4917bcd9d00970bae6a7b33ee9287c05

Observation ba7dca3b-95c3-4ad4-bb5f-c690046d5165 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Gpqa: A graduate-level google-proof q&a benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.595871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.595871Z digest=sha256:56e0a16b8b794200bafd517c267586d9812f7bd5f28a691d2a7a8d74c7a9d743

Observation 4cb57203-05da-4b9f-b0ed-a5a09edc8c99 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Code Llama: Open Foundation Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.600900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.600900Z digest=sha256:f4242e388c4bf09f26a4320c44b55da9b88160a0f92b0926657f55e3a2a3a84a

Observation 7b37e65d-e30c-4cab-b8ed-574fb9b42aea · outbound

This paper cites A practitioners’ guide to transfer learning for text classification using convolutional neural networks.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once A practitioners’ guide to transfer learning for text classification using convolutional neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.777088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.606428Z digest=sha256:3438651aa71d4fa5033e4e2f5dbcdac050b81688159440606e0bdb428dcad3de

Observation 1b4d177f-3de1-4cb5-a6d6-90521a19064a · outbound

This paper cites Rethinking Reflection in Pre-Training.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Rethinking Reflection in Pre-Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.611579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.611579Z digest=sha256:7af6661ac946cd4711a048681b9095cc00081e0c8985c3f9969c404cb8a0bcd5

Observation 6f9f2f83-dde7-48ee-8ea0-60bc1d38bb34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.617129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.617129Z digest=sha256:419d555e685d4653a8bbe7924f1455e1394594166943eb73e468490a2bb046e9

Observation 2681144d-be27-4ec2-9346-bf4175310e8e · outbound

This paper cites StructuredRAG: JSON Response Formatting with Large Language Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once StructuredRAG: JSON Response Formatting with Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.623431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.623431Z digest=sha256:62d2720865706c5870a51f102a6a02985c4fedb05683c919f906acde4a84d26a

Observation 1f1595e1-a2cb-4814-bf27-bb3bea871c58 · outbound

This paper cites an unresolved cited work.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:33.734942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.629984Z digest=sha256:764ecd47422027f9b618dbcd7dac6ac5a74b4209204d39ed7bcab148f11e2079

Observation 44f19072-1194-48ea-8190-4c249b62aa5a · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.635546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.635546Z digest=sha256:f314baee41efe97f981ac6086406a5f78991b50c7efd0addba50f2ea6fdb4332

Observation bca4952e-b9f3-4f99-b3da-19b652ad3bb7 · outbound

This paper cites Commonsenseqa: A question answer- ing challenge targeting commonsense knowledge.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Commonsenseqa: A question answer- ing challenge targeting commonsense knowledge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.702329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.641299Z digest=sha256:0305f4698b43b09964dcd50e17a09667bfe5944f888d77d66e7398885cb33cab

Observation 3592653e-52be-4764-ac17-dbde88ab955f · outbound

This paper cites Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.645965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.645965Z digest=sha256:17abe57aa559062ee907ee033cbf40eba8a8838234d2c1923130b293d12bd3f9

Observation 4c333033-5f40-49a3-8bdb-7d6c65dfd001 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.651233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.651233Z digest=sha256:19e9d5b03e2c300271e63d3037998e25927bf1414a3ed3705a45d3fa367b311b

Observation 04ed657f-18a9-4afa-884e-b8c9effaf7d8 · outbound

This paper cites Gemma 3 Technical Report.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Gemma 3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.656844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.656844Z digest=sha256:fc388ad3945690a8fc5301bfc8b996ca2782dcfd4650eb0cff0ba802a8adb1a2

Observation f127b587-d06f-4894-8541-986dae1a3acb · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.662296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.662296Z digest=sha256:18752f6018aa6c3a91a7b0f163996b7fd56dc4487d0911c0a09e6714ddcde60b

Observation 9f0747c9-feec-49a6-a2ff-1d24a093019d · outbound

This paper cites Open Thoughts, January 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Open Thoughts, January 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.670553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.668855Z digest=sha256:8361a9c314d07d124a3e09cfd02a7841905af0b4429b10d7be9f7ee64bc29f99

Observation 7e17e468-f86e-45dc-8610-5634a5272ccb · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.674386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.674386Z digest=sha256:2792d8f8efcea81cdc7e683c85422efdf8e583ae364065239495ae502356a680

Observation df5bbf48-696d-48e1-ad0a-e5f70cfe8bd2 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.681214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.681214Z digest=sha256:631df2de8e0f0e57b8d9bf2511aa9f69fbaa89a8f7c1f7bd711151f82bdc8ac0

Observation 735fa473-68b5-4aae-b2d0-b8187d8135bd · outbound

This paper cites Evaluating llms with multiple problems at once: A new paradigm for probing llm capabilities.arXiv e-prints, pages arXiv–2406, 2024.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Evaluating llms with multiple problems at once: A new paradigm for probing llm capabilities.arXiv e-prints, pages arXiv–2406, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.631351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.686947Z digest=sha256:0e0a7689cac397212131b5286b51946f6acbae9f9f7f79766a35b7178a3dd921

Observation b60c03f1-14c4-4a95-a520-91ce8f4fccf3 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.692436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.692436Z digest=sha256:21495029930969e3b051fe14ea983952951f681887ae45236206175e33f89827

Observation 5775c902-e633-4850-8d68-933baa18105b · outbound

This paper cites Qwen2.5 Technical Report.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.698122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.698122Z digest=sha256:e3c56ff103cbf7d6b8dab97b328a6099ce03fb37a51708aeb14817c0db875604

Observation ef69f0ad-aad0-4073-b79f-7f9ea4ac61cf · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.712672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.712672Z digest=sha256:09149ed95e3d6a1bfbbb2bdfcebb5293d6e7c553514e4174bbe5affc8fb8ee49

Observation 111b102f-a05e-4ae1-bf88-8700d95d4e91 · outbound

This paper cites Aime-preview: A rigorous and immediate evalua- tion framework for advanced mathematical reasoning.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Aime-preview: A rigorous and immediate evalua- tion framework for advanced mathematical reasoning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.605064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.718497Z digest=sha256:9ba7dba8a690866529dfcd59aaf48432ebb650d3b4b2329db332eabfbfe71d94

Observation a62d3ed8-d201-482f-a455-a323b0323cab · outbound

This paper cites Mitigate position bias in large language models via scaling a single dimension.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Mitigate position bias in large language models via scaling a single dimension

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.581735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.724414Z digest=sha256:66bd0675ade11c13070ef20829e2598e824d84f3be6a336e4b8be7b7f5452473

Observation 1b4084d5-1f64-4f1e-97a8-33d14e2df46d · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.731325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.731325Z digest=sha256:3a41c2a907c8c6e6080416f21ed3a4225d8646fba5f3f8670b36b7b31feff6b7

Observation 79d7c9bc-b709-45ee-80c2-f5b89985d4fb · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.738286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.738286Z digest=sha256:1bdb6b4d5ded7d67b910bff1c9b5a4c5fb17983ea071eb61cc3d94a5689a8ee3

Observation 20b70d40-304a-41d7-be02-63be69b72be2 · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.748303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.748303Z digest=sha256:de13b8c6edce122519f2b061ed7da581adcfeee9eceeb6aee9e90eb535949ad6

Observation 77c51be8-50fa-40f3-bd10-77489d4565de · outbound

This paper cites Your task is to extract the final answer from the prediction as it is, even if it is incorrect.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Your task is to extract the final answer from the prediction as it is, even if it is incorrect

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.556386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.759917Z digest=sha256:8f09aab8f40fc29b942251248939c1849144b3096fe954772a7a424c4ad8445a

Observation 331cc45b-32e9-4aa0-8c52-0f616cd45472 · outbound

This paper cites an unresolved cited work.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:33.530155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.767438Z digest=sha256:b5b83b465ce01f5dab8f64da364ebe35da69ee3b283e10d8dae7326902d36531

Observation 4c5a31fd-82c9-4ed0-b29b-3b8b337a5bbe · outbound

This paper cites You should set the final answer to None (e.g., \boxed{None}).

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once You should set the final answer to None (e.g., \boxed{None})

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.496129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.779302Z digest=sha256:9272ae39b0289ec107eccda70d414082aa6f807f72f63b77f955a1884c239b0b

Observation ffbd914f-2835-49ec-aa56-1a60046abf66 · outbound

This paper cites an unresolved cited work.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:33.471313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.785220Z digest=sha256:13128ab31c93c5f1daea07a9b4498f68c6d90bf5127529289a5541aae607bdf2

Observation e76c610e-89fa-43b8-bbc6-310aca6f4d4f · outbound

This paper cites For example, if there are three questions, the output should be Answer to Q1: \boxed{answer 1} Answer to Q2: \boxed{answer 2} Answer to Q3: \boxed{answer 3}.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once For example, if there are three questions, the output should be Answer to Q1: \boxed{answer 1} Answer to Q2: \boxed{answer 2} Answer to Q3: \boxed{answer 3}

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:35:33.441592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.790814Z digest=sha256:2241651e05e54bf4f94bd215dbd4d375db3ee876bbc059629f19d9ff206d2f7c

Observation eb0aa1e9-157e-4a91-bf2b-81c7ca3f7619 · outbound

This paper cites We need to find this distance, express it in a specific form, and then compute m+n+p where the distance is m√n/p.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once We need to find this distance, express it in a specific form, and then compute m+n+p where the distance is m√n/p

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.400304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.802279Z digest=sha256:18c375929153163d0f24892883da5ac35f53430ddd3d6a2b6b62131751df86b4

Observation b69446d0-8b2e-4d0e-855c-f6703ce1eb05 · outbound

This paper cites This distance can be written in the form m√n p , where m, n, and p are positive integers, m and p are relatively prime, andnis not divisible by the square of any prime.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once This distance can be written in the form m√n p , where m, n, and p are positive integers, m and p are relatively prime, andnis not divisible by the square of any prime

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.421756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.796838Z digest=sha256:34b479e7e059917b0265e669f80e80ea8816031cc85f7b05ea752bc506ac6ed1

Observation 9b5446c9-9f63-4f74-a0bd-2388734482a5 · outbound

This paper cites tikz\"); label(\.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once tikz\"); label(\

Reference 307

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:35:33.371631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:35:31.807793Z digest=sha256:ea2e4d25bb69a49375a162424702f223263ddfef18fecf91fbe537152d74f095

Pith citing papers

Observation b4698c59-ecbb-4409-b14c-54f31f80b53c · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.542163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.542163Z digest=sha256:2d8dd5f8f4592eb6966c732dfb301be6bc73c553385b447a08c3f6c96c0e8c42

Observation c7a683dd-5643-4caf-aeaf-4f7f986a0c69 · inbound

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure cites this paper.

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:54.185654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:43:54.185654Z digest=sha256:89edebebeb8ca8be44e2b180850c7348812437167d747daee7bec61f0ff68ece

Observation 809f36c3-16c5-4fd6-8885-1d3af49c5462 · inbound

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs cites this paper.

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:06:00.571138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:14:38.371700Z digest=sha256:5c35a0782bfc4a6e28be10fe3d7786231581ce565ce4c795817ae31f5362b3f8

Observation 33210c55-ec7d-49c0-8eb3-8996eac8a493 · inbound

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents cites this paper.

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T14:26:20.341497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-09T14:16:17.970333Z digest=sha256:b488ad021227f23e0605ecf8599fa449f6bddc6544e051c4fee17b396422c133

Observation c6db5c40-fc03-4847-ac89-107798b2046b · inbound

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions cites this paper.

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:41:01.511851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:41:01.511851Z digest=sha256:0394c5f54ad26a89239040c919a954f7be67d44fecd52a7fd7da9b19fa61c5f7