Pith. sign in

Paper Citation Record · LEDGER

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 14 inbound Pith citation observations for arXiv:2502.05252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05252 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:16.984092Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:41.438347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:19:23.876486Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4bb915a-8c43-4bb3-93a5-c825f3db962c · outbound

This paper cites @esa (Ref.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? @esa (Ref

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.740385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.740385Z digest=sha256:b0fce20546c3aea4ffac64df8bc8e1623bf4c1a260bf3f7ca79b0480b2a679f9

Observation 8ee7e233-7850-4e89-815c-7d80b9ef8b8d · outbound

This paper cites an unresolved cited work.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.747628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.747628Z digest=sha256:67b0760821dd1f212a1d09f7910e9e811527e6d669d11d41a486a795b63fdd1a

Observation 23794d80-f9ca-4366-8a56-22813ae5adce · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.753303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.753303Z digest=sha256:4ec9fb0de5ffc5190e7efafb4dd99614766a2f33aaa5fe4047eb5de75372b6f2

Observation 767b9b18-8209-4b7e-b845-7ff8bba35e4a · outbound

This paper cites write newline.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? write newline

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.761081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.761081Z digest=sha256:1c716af9e95d616a44fbbcc90c0b67b6744a547a5ce1a15fb1f9353f872b5883

Observation 70001e99-033b-4608-9428-9cf998ad5983 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.766880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.766880Z digest=sha256:88b85fc33b94b04df4d9d0dffccd256ce2d55f143c97ab2fdcf3fbed01842708

Observation e456d5dd-cf02-4266-848d-c477b5ed4276 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.772448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.772448Z digest=sha256:3afaaf21d47b7e8b591e92afb4c2574b3f71180fb2b6d391f512432dc99d5ed3

Observation 37095ade-0bf5-4673-941a-15fc4536eb9f · outbound

This paper cites Longformer: The Long-Document Transformer.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.777759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.777759Z digest=sha256:91602e0ac9714562b6ee9315f24ee3667a56de17e153ffbd58635b5032426a6f

Observation af607390-b8c6-49c2-94ef-527c11dcf5ae · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.788920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.788920Z digest=sha256:9ef86c29970dfacb7e54e65857dea8de3d1462fcacd9bddb5727ba378616800f

Observation 2816ea9e-b8a9-4e7c-a8bb-2b4aed73cd54 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.800528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.800528Z digest=sha256:aae53d2a316169b46258bd449eea22a7feb39220c2e7011bdffde639a7b3ab96

Observation deb7c6a1-35a5-4fd0-98f0-f325b302d22c · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.805549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.805549Z digest=sha256:89103cb7aa2e3ff54d13d12cd64fce55444009dffb93a7188a5c19eb3d840d31

Observation f22df550-9487-4c3f-96d5-4cc61017e205 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.682833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.810585Z digest=sha256:68ba0268b10b6747b1f7496de45a9c0eb1286630a2f7179d9961e56d41dcb781

Observation fd4ba3bb-527f-403a-9c2e-188f9529c300 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.815476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.815476Z digest=sha256:ed97b2523dccc8284768e8ef1578717f36dce770346442cd60c01a8cccd139be

Observation cc7a17b9-e075-4884-95d5-0a71615f0d35 · outbound

This paper cites Neural Networks and the Chomsky Hierarchy.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Neural Networks and the Chomsky Hierarchy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.820859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.820859Z digest=sha256:c9ed7151eda425c137f65bf28d1d751ff51a72e4c6b59ea800131219670c3a9c

Observation c2fa464d-b6be-449f-96e3-ecf758f6a627 · outbound

This paper cites The Llama 3 Herd of Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.826047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.826047Z digest=sha256:154db7ad6bec77b2b30d3b7a73ce820c9ad1a2a4ac7b3c103588a68bf96fedd4

Observation 911d184e-580f-41f9-9fee-f9ec652b0794 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Needle in a haystack - pressure testing llms, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.664012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.832004Z digest=sha256:43dc7e717562e7c084b0e8b97bde63e7a4b2c63a2e89541ae058168bc472806e

Observation f57ef75e-9fff-47cb-91c9-5ab88eb228c1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.837781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.837781Z digest=sha256:19867c5d1c291bb6152ed07f8cd7359ae4aaca42115ae1b76cb6b44665f1cffe

Observation 2a8d70f9-c571-4ef6-9661-5c72e8cb49e8 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.848492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.848492Z digest=sha256:f1748285da2c47cd0f81e4fc1aec2bcf8e6f874b6a8bab2631c88aae605b7005

Observation 8b667d32-78fd-4e74-990b-52721a025d7e · outbound

This paper cites Mistral 7B.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.853620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.853620Z digest=sha256:74c0e78106df1315489eed173cdc79ba547b3d9f73611edd669f35dd9eac9ac2

Observation 5cb60de0-520f-42c2-91d2-faa9b7e76904 · outbound

This paper cites Active Retrieval Augmented Generation.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Active Retrieval Augmented Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.859507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.859507Z digest=sha256:635f771cabc71723b87eae8becd8e5df3ebf7076bae2e2a5b2522c685dcdf23c

Observation fa549915-4c4e-4630-b053-f062f4cad42d · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Needle in a haystack - pressure testing llms, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.645657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.864472Z digest=sha256:889ab24768b33906dec98abe75086d6d75eadc1de4ad1df936eb970f012eddbf

Observation 190ef306-8a5d-4f36-ba38-65563c3d4242 · outbound

This paper cites BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.869442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.869442Z digest=sha256:ffc808238fb27247ed64a8467361d9c4c95d439851f83cef933a2b7c14619ece

Observation eb849715-65bf-4fe6-9e3c-4df0bb4bb99e · outbound

This paper cites Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.874664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.874664Z digest=sha256:7bfde1146a21c2fa5083eeb21637d214301cf5968be275c5b5b84960e780c270

Observation d5b394f4-c6a7-4a30-8bc6-ee4cdde6c0cb · outbound

This paper cites Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.880139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.880139Z digest=sha256:3f6007a6af908337250177026b3879374769eb16f0f3e3e45fcde3988d0e7d71

Observation 2a7a3af3-51a2-46c5-bbf0-96150d6fcedc · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.885207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.885207Z digest=sha256:73420062855937d41f312472062d91b01d0434e9d4e280b08a515943ba28e270

Observation 937d0ff0-3e26-4cde-aab8-2b32467d4ba9 · outbound

This paper cites Long Context vs. RAG for LLMs: An Evaluation and Revisits.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Long Context vs. RAG for LLMs: An Evaluation and Revisits

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.890221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.890221Z digest=sha256:3d71a157df911e1b41ed4c5e0b3a108168c879213690258da133341d44e290ab

Observation 2f5691aa-e510-4aca-a989-4e9e6de23538 · outbound

This paper cites Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.627462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.895307Z digest=sha256:d30c23439c9de717f455266bd437a803b5449bab2c2235fb875eaf1c63e174d5

Observation 4520cab5-8190-40ec-9c3b-db670fd2140e · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.900015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.900015Z digest=sha256:973eedeac5b28f769bddd375ed07d52038cf514a8dcb5d5624e138019a3de35d

Observation 1292359e-02cf-47e3-826f-2e24f6be073d · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Lost in the Middle: How Language Models Use Long Contexts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.905007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.905007Z digest=sha256:1a14275863b76607f9e675b13c46e4c9c8d456abea5be2258800ac7a85c1ab77

Observation da20d73c-593a-4b08-8a0b-38a07b2b10af · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? DafnyBench: A Benchmark for Formal Software Verification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.910497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.910497Z digest=sha256:40313c3c0422b5f9709fcb33465a4787e1cf43e5c01a4e69219ebddef755a409

Observation e4214ef1-4b51-46d3-9c0c-f24b46781a31 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.916090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.916090Z digest=sha256:3e1cea2ca1d6a4a0561e55d7a0ef9a3a596fa283f6b03f97da34bb336dbfe7d1

Observation 34d51ff1-5307-4ae5-b21c-113fca06d93d · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.922469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.922469Z digest=sha256:cbbf2f0aff1dc064cbaa4f9fab6b6928ae4e3b56ce7acec6bdb113a0aabc7230

Observation 9947327d-b5e3-4c06-8b34-13b21beb5bf3 · outbound

This paper cites Qwen2.5 Technical Report.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.928032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.928032Z digest=sha256:27e040d3f7f17bd06a4ec540ae9467da3b09909c6d36ec7b1eba56d5afe26f8e

Observation 516599cc-a18b-4080-ae52-665b23f30ef4 · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.933498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.933498Z digest=sha256:f25de7d8fa58fba553cf0178b7c23b9a48c1ec07a1a98b51a6bea3cdc8104c9d

Observation 208e2cbb-bcb8-44c1-bc73-2e1eead2fe66 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.939267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.939267Z digest=sha256:dfe6d5c2441a16903839416ac0d1c8bcc0cb5d1f8c7fd08b02fef7dde8d359ba

Observation 2f2b01ab-d196-479a-b1a9-a493a486b12d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.944984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.944984Z digest=sha256:150ea15887bede974d71d6b3ee6ecf7594627add51abec9db2d9bd17c2f87f71

Observation 82a51e3a-6c9a-418c-8cad-62795881d5b4 · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.951317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.951317Z digest=sha256:c55cebfdf43c5089802af3ffefa7c14f3672360f253bb9a88b5f47465ea82f7f

Observation ee703469-e33e-4977-bfe5-8e082c6167ae · outbound

This paper cites Modular elliptic curves and fermat's last theorem.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Modular elliptic curves and fermat's last theorem

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.607210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.957513Z digest=sha256:6ad5f5c84b496adc53cb06baa4e75cfc25f70f2ef2ac39bb17e716fbef1b3063

Observation f8675a3e-e296-438c-86b1-91e4ee6e5ac3 · outbound

This paper cites Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.962705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.962705Z digest=sha256:fdcacaf9babef1f90d0f025ff17fa3dbdcba453ad2f6b0482af31a9a7f52cbdb

Observation 542152b1-4ed4-4c88-b19c-2b68971428c2 · outbound

This paper cites Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.968703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.968703Z digest=sha256:00044a3aa6c6bf03c8d600166a304342b4c0ba6d5c71e18a10865e44a16fdf89

Observation 87b9c447-6dd1-4c54-ac8d-e1a7496cf3d4 · outbound

This paper cites In Defense of RAG in the Era of Long-Context Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? In Defense of RAG in the Era of Long-Context Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.979011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.979011Z digest=sha256:d261b2853810c892574b92449d95b7bb5b6bb9efa9da33f98ad48e4aa523df56

Observation c4e73476-9160-44b4-97d4-1b5ad729b0da · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.984092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.984092Z digest=sha256:d4d81ad22c9ffb94043705fc970afd2108eddd3b2881a778988cccbb57a38ebb

Pith citing papers

Observation bfdbaa7f-bb57-45e9-a4bb-78d238f20d31 · inbound

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings cites this paper.

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:41.438347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:41.438347Z digest=sha256:91963d19f6b4843ff3999b694d94ff81a79ce4d9b03abb6df1b492d439c3739d

Observation 2448123f-d8e5-4f46-ada2-a791e302adcd · inbound

LongReasonArena: A Long Reasoning Benchmark for Large Language Models cites this paper.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.176410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.176410Z digest=sha256:597ddc6545cdeb3fd2c4dc270ee5e1a813a0e6d7da6a57e01082da9d05cf3b8f

Observation 9c9a0ee7-7f47-4758-bf05-0b9669aa0dc6 · inbound

vAttention: Verified Sparse Attention cites this paper.

vAttention: Verified Sparse Attention GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:10.093674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:21:10.093674Z digest=sha256:ef6800980d0757e7d6658ebf8f7fd92794ce842638afbe7e35e5ed6e582803e0

Observation 35219edd-3f3e-48a5-8495-413cc13ba9d3 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:33:32.819008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:0d467fed0c148d487a12c3cae49ab0978abd095a69df6107237bb350b8a13050

Observation 513603ef-f1ea-4074-be2c-07e15eadb755 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.602502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:0e097bfc72ef8fcf070cd1e3a186f355d4b262a9e77937dcbb2b4e58129df533

Observation 0ab7533e-c155-4bc4-91d7-dcefc186051d · inbound

From Local Corrections to Generalized Skills: Improving Neuro-Symbolic Policies with MEMO cites this paper.

From Local Corrections to Generalized Skills: Improving Neuro-Symbolic Policies with MEMO GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T15:10:27.013215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T15:10:27.013215Z digest=sha256:2ed023ba4361d0606a1e6c48854c2b5dd4c072b219a88ac85a1e315527003c42

Observation c81149b1-e343-401f-9449-19308cc205c0 · inbound

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning cites this paper.

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.740864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:34:14.413131Z digest=sha256:365a032fb76a1180d6260f3e8895d66b3b04cd27c1e1cb606cbd04e877c46a31

Observation 74e4b96d-0ca5-4a0e-acd1-7008e2a33d6b · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.190972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:07a5b42041c9f5c04547965cddfce5ff53dc7c02116a5c0f8cf311eae61e36bb

Observation 1d00e09b-b71d-4e6e-8d9b-49d4d7b10ae9 · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:2794a686bb5f08ad50daa00efb9974c05c00a3e6131f76b07229b9d40a31cdd2

Observation bb9b8bf6-909d-405b-aa7d-9a7e2cb7a8e6 · inbound

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding cites this paper.

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.770407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:25:20.511933Z digest=sha256:df31c08cebdd440fea182b2750aeb35e8294114f8b806f3eccf89c8e3bfb15dc

Observation 6a15e411-70c4-41e4-aa92-dc30f06307ed · inbound

ATLAS: All-round Testing of Long-context Abilities across Scales cites this paper.

ATLAS: All-round Testing of Long-context Abilities across Scales GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.979086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T13:30:09.778806Z digest=sha256:ac7c8765ecc361097844583e7069d44c9a19f34b67572b5e1ef1b1b723ac7687

Observation ff892fdb-d0b3-4522-8990-5e2477da3a48 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.878014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:4f2f67bc98eddb025c331f3db75c81774ec526b50ee4bf289c1c9fc87152603f

Observation 09806160-cf7e-4d87-899c-67888cbaa6d6 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:04.371774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:04.371774Z digest=sha256:f9aa5f66579b55db9b6d9c919307bd13600ad6447113f4c7074fbe948b0710a1

Observation 3a8dba01-3ae5-4436-b557-10af259cd432 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.900522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.900522Z digest=sha256:3b4415978c297ba84c3e0bf6cbb63a35b5c9719cf1f3a40eb9654f21a388793c