Pith. sign in

Paper Citation Record · LEDGER

Scaling Reasoning without Attention

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2505.22425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22425 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:28.835572Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:18:54.163862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.879253Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e310493-4137-4ef4-ae9d-4c5f6eff8019 · outbound

This paper cites GPT-4 Technical Report.

Scaling Reasoning without Attention GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.227804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.227804Z digest=sha256:136757d8703c00e106c40fcd5ca4fc833db0fa857709796d99ac3d6b29195987

Observation 8cc3dbb4-c48c-4a46-a9f4-ec87798285cb · outbound

This paper cites Language models are few-shot learners.

Scaling Reasoning without Attention Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.415360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.415360Z digest=sha256:56af253d15ff864ec279ce1e4294e1c90755c981f6950eecf8dea33dba5d02c5

Observation 9fc8e080-b323-42ca-86af-59ed924e7fd5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Reasoning without Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.648195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.648195Z digest=sha256:9d9fb20b3a8af344d34ed9be196c39decb320020f374d7bbabaa8f2e0337445e

Observation 69b356c1-0c92-4468-8a44-12d431ded2d3 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Scaling Reasoning without Attention OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.740105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.740105Z digest=sha256:2a068c89b5d90fbc01b90bdab947d7894959368976351fcb641e3a5bf7e166ee

Observation fd5febec-8819-4651-ae3e-7af956b07aed · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Reasoning without Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.815329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.815329Z digest=sha256:28eaefef67ffccb2618ad50564b16608c838604c326009f83d3166c5227d68b2

Observation 709c1181-f750-43e5-a30e-b624d8459e01 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Scaling Reasoning without Attention LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.947907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.947907Z digest=sha256:f8b806ccc6274f173a3382d65183c894b4d488ef34312f6d806aa0dd4374aa42

Observation c66d47db-93db-46ff-988f-a33b0611c00b · outbound

This paper cites 9 Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ra- masesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al.

Scaling Reasoning without Attention 9 Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ra- masesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:30.398256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:13:27.038519Z digest=sha256:42d14333baa72d19c6f18c651bf71c26f4548e335c1c63099baee1d022a19aa1

Observation 1bab0e03-c8d6-4045-b22f-7be5514dbfe3 · outbound

This paper cites Let's Verify Step by Step.

Scaling Reasoning without Attention Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.137705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.137705Z digest=sha256:804e44c91f00c059dcccbc97fe07a2a62ca2987a994b131d38f5c2ba7bdcc5af

Observation 64cbe68b-78ec-4b6c-a9c0-6af0de2e2f80 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Scaling Reasoning without Attention AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.227849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.227849Z digest=sha256:f2a7ab97a836cd210958f5f57dc9f2d7d21556475ea76f6fe0a8842732924a82

Observation c9926482-f297-4aa8-b329-8cf8c955e4f4 · outbound

This paper cites s1: Simple test-time scaling.

Scaling Reasoning without Attention s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.340471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.340471Z digest=sha256:e6709e3dbe63169e1391c31e2262dc903e7eae9ae3f0b5ae8582c3d3db73277a

Observation ab333d8c-9054-4dee-be5f-ae0b0743df95 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Reasoning without Attention RWKV: Reinventing RNNs for the Transformer Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.469238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.469238Z digest=sha256:ff2f5024b7ed54b41fe0c52121cd2558ed50bce5a41bb3a95cb3021ef7424a5e

Observation bb416504-5ae1-4418-bd1e-c5e554c7af25 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Scaling Reasoning without Attention Retentive Network: A Successor to Transformer for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.624874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.624874Z digest=sha256:e5b55432886bcda6a207d7992bb6c3ef8799d00b72b7d7837332bab745196b42

Observation 5d3c91fc-0bb1-4550-a45f-cd0a3c0f3b01 · outbound

This paper cites Gemma 3 Technical Report.

Scaling Reasoning without Attention Gemma 3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.735999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.735999Z digest=sha256:8af056eb2e7298a652da7f80a4d745b60c337c1b4cac43e6f95aa6ac6b71b03b

Observation 6d6a2ed3-9257-4b80-8998-86bac1a0fbbb · outbound

This paper cites M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models.

Scaling Reasoning without Attention M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.861259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.861259Z digest=sha256:976697ba0543d1fa08ee30e2bdfef4ef0f0e566f4b3b9450bc4f0c28b53e4c3f

Observation 11a2ed28-ce7e-4a00-9c4d-92da4a5d57b2 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Scaling Reasoning without Attention Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.970162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.970162Z digest=sha256:504fbb3670877256f7bd18c9d6e05f1f75db046e5c1ac90325c17233ff0e1008

Observation 232cce74-e538-4a71-8cb9-6ad92e78a480 · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

Scaling Reasoning without Attention Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.091568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.091568Z digest=sha256:122248327e7963bbeebf97fbd941aa934e55bbc75858bdb50ea4bff7dd43933b

Observation 49190441-b27f-4683-bd0b-ae4a2a8fc160 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Scaling Reasoning without Attention MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.195543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.195543Z digest=sha256:423782c1b34c71b50a04c6ebee34cb7d70ebdd12a7339d0421f73cda56d70301

Observation cf98c9ff-b42a-496e-9dde-b54d27aff825 · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Scaling Reasoning without Attention MAmmoTH2: Scaling Instructions from the Web

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.299855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.299855Z digest=sha256:9e6e665cafd18fc1c87d2ee940bd49b3816c6a27dc59affa60ceb4a9701cede7

Observation a25caafe-4ae7-4836-a253-4895d8bf899e · outbound

This paper cites SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving.

Scaling Reasoning without Attention SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:13:29.309591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:13:28.389604Z digest=sha256:7191b82a327dd48a17a35dd6d0d6c994ecc39e833b7774dae5f59b23b282f7d0

Observation fb20beab-3a9c-40c7-b91c-7a9042483b9c · outbound

This paper cites SubgoalXL: Subgoal-based Expert Learning for Theorem Proving.

Scaling Reasoning without Attention SubgoalXL: Subgoal-based Expert Learning for Theorem Proving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.536244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.536244Z digest=sha256:2f0e40e1d2dbfaeb9d49ce6d1b502cb6de0e31e34958c0cd828b700e6880961b

Observation de3232c6-d290-402a-ba76-56aadcff66ee · outbound

This paper cites Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models.

Scaling Reasoning without Attention Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.644116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.644116Z digest=sha256:390ac4d3fcfd61369ef682075110527c19edaf73083ccfd89a0e911b4cf4564f

Observation 8832bc9f-8306-48f5-bb41-1d515f79f604 · outbound

This paper cites Efficient Attention via Control Variates.

Scaling Reasoning without Attention Efficient Attention via Control Variates

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.758966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.758966Z digest=sha256:afc9c04ded6547c05765d392ea978421d1dd529444d1255bb23e86f8b197d26c

Observation 069d62ff-3531-4f02-a1b6-9827a79ec468 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Scaling Reasoning without Attention Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.835572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.835572Z digest=sha256:af8ee17fe9026c6359a33f950303f9a8403765c45c4e1deb8ed44390175e98a5

Observation 946022f1-e3eb-409b-a108-8f113b4e3add · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Reasoning without Attention Evaluating Large Language Models Trained on Code

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.496858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.496858Z digest=sha256:fc42f37d7c809c73b9d0e2a3df8dd7db577e0b3ba803ea60f1217f489ce0bb1f

Observation 377417d6-7b14-47eb-95d5-5826c4681d22 · outbound

This paper cites OpenAI o1 System Card.

Scaling Reasoning without Attention OpenAI o1 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.876946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.876946Z digest=sha256:06459b9430891e5080fc9c91d962ad85cd9d0f937c837006441e549b271275b3

Observation 0379efc3-ba6b-4810-8651-41a7ca3c7fc3 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Scaling Reasoning without Attention Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.038685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.038685Z digest=sha256:98aaf00bfb33465c5f4fb67cabb652a039b91370195c85883506c5d50768b688

Observation f7c5e5c4-6f0a-43ba-9152-fe1d4437403c · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

Scaling Reasoning without Attention OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.292420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.292420Z digest=sha256:1bb81de61d253ce0a398642e981a8f14178c8291c6fcc3da84e60164e7fdfbad

Observation 48ba9ba7-0e97-42a7-afe2-ebc9327e91a6 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scaling Reasoning without Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.570912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.570912Z digest=sha256:0885b10766284a37cb55d225ae9cccb3886d85eb6ec492e4b7d0b3a285a4461a

Observation ea777f1a-5425-4c36-a710-89fe317673d3 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

Scaling Reasoning without Attention Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.357993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.357993Z digest=sha256:860fdc1498b24ff6d4b38bd9b57c517e271ae78f55247beae96f5f3b8dff43b7

Pith citing papers

Observation 38fd6968-c2b3-4337-9c67-ebb554f29524 · inbound

MetaLint: Easy-to-Hard Generalization for Code Linting cites this paper.

MetaLint: Easy-to-Hard Generalization for Code Linting Scaling Reasoning without Attention

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.876223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T04:07:31.283348Z digest=sha256:500ff04ec4d4da0ec2fe87f46f1508891e2d6d7b9a3726cd277f59f954dcc9e6

Observation 176b3fa8-1938-42be-a182-b772f314f799 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Scaling Reasoning without Attention

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.880509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:c2c9e87d0cc3be929aca0a9b125caa00741c9e25597635f5c0166f7b73ffc652