Pith. sign in

Paper Citation Record · LEDGER

Efficient Pretraining Length Scaling

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2504.14992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14992 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.825037Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:31:30.334937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T07:43:11.766620Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved50
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a495a46-b68e-4c72-b08b-8ed75d565a47 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Efficient Pretraining Length Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.558018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.558018Z digest=sha256:6f98cfc8ebfd0237d436131c0a83ad92d8b453cb4a0b8a83bf07ab39949919c4

Observation 37f726f6-5bd4-452b-b1ee-51c5db3e5e8e · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Efficient Pretraining Length Scaling Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.563405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.563405Z digest=sha256:65080caf4d7543788050b6cd59c67e3b0c1860eb3436c88e2ce9483eb0bd2b8e

Observation 15b54ffb-3f76-4735-b1f1-a3cc8cfe9928 · outbound

This paper cites Longformer: The Long-Document Transformer.

Efficient Pretraining Length Scaling Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.567873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.567873Z digest=sha256:ee8d68305564f4be563e008370724fc26095f4c9f3b0efb91d7988bc6884c0c3

Observation 36e29e45-f8e2-4088-83a9-6c5861deb028 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Efficient Pretraining Length Scaling Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.572503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.572503Z digest=sha256:94ca4fcb1b2e47e493ccaf37fb378c793d70716b75250d62607ca31b74825b60

Observation 091c1abf-aee0-4067-8db9-777408e2aedd · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

Efficient Pretraining Length Scaling Striped Attention: Faster Ring Attention for Causal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.576790Z digest=sha256:3a5f6c3afe39408fa7b1a735248b29e9892277c48eab5f042e36c2f319596e6a

Observation f0f1e1f3-92a8-43a5-b4ed-aa5dc82d590b · outbound

This paper cites Step-level Value Preference Optimization for Mathematical Reasoning.

Efficient Pretraining Length Scaling Step-level Value Preference Optimization for Mathematical Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.581679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.581679Z digest=sha256:f1342387a395b1fff0c31fc88cc87430f4358285851123865e03a679e3dc8b3c

Observation a9ecf90b-d3f6-4c55-ae68-2d61dd38abc3 · outbound

This paper cites Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking.

Efficient Pretraining Length Scaling Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.586615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.586615Z digest=sha256:fdbf68ee735962122f0cbd2ec5439f00e1abb64ea8867db296335c2dc42d24e7

Observation d98a2357-b202-48e5-ad30-4a0bf0bd91ea · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Efficient Pretraining Length Scaling Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.591194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.591194Z digest=sha256:57ec7c28ee31e953f336fdd0c5f17d86468a1b49b3cfdee6ec785cf731ce216b

Observation b9e293a4-de22-4678-9395-6af6dcc67457 · outbound

This paper cites Unified scaling laws for routed language models.

Efficient Pretraining Length Scaling Unified scaling laws for routed language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.595724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.595724Z digest=sha256:c8fb0bfd94e0806837bc340057d9c469cdca0b48a070caacd2d906b864da5f63

Observation 80048c91-2a95-4320-867c-e50c8cc87163 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Efficient Pretraining Length Scaling Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.599886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.599886Z digest=sha256:1912719948e0d78078ea5540ace3cc91b95f8f791936993041bed2194d5fde5a

Observation 665c3ae4-19f4-4b9f-bc13-659ac7dbb42b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Efficient Pretraining Length Scaling Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.604174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.604174Z digest=sha256:d526522113c8ffc0f8e23cc225313bbe3178bc7a098bb9e206ecc099917e2e2c

Observation 8a8c855b-09c9-45d8-a140-9ace5b336d3f · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Efficient Pretraining Length Scaling Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.816535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.608654Z digest=sha256:2fbb4cc821a0126e1a2ec2591adbb398228926f9a6fb2d9c00260f1c89df7589

Observation c858dc2e-bee8-45ca-b2d5-c2fbc0795e09 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Efficient Pretraining Length Scaling Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.803359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.612738Z digest=sha256:95a39daae8fa1ab933d36bd8e73c577e8cbf96a2376840964fb7643ea8281279

Observation a35e5cf4-34aa-4230-be4f-26896c2c47d0 · outbound

This paper cites Flash-decoding for long-context inference, October.

Efficient Pretraining Length Scaling Flash-decoding for long-context inference, October

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.789928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.616727Z digest=sha256:2ba14ec18567b347c5c17cb0e84ddc5120d6a996ea8d243bdddee4a1b3139a8e

Observation 0fcddfd9-b261-4dee-aa43-098683fe5ded · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

Efficient Pretraining Length Scaling LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.624969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.624969Z digest=sha256:e2fc108b89c9280bc0da109aa03cc17159ecffe6604364148e6cdbe504b485a2

Observation 90fe73a9-05d6-4945-bf02-11a2bcb339ef · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Efficient Pretraining Length Scaling Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.629386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.629386Z digest=sha256:3841b49089fde8514e550b6a50646807acb2aa1cf7ffbf99bdb6ac0baea3bd6a

Observation c26e4082-807b-40ae-a037-b99d19b14464 · outbound

This paper cites Think before you speak: Training language models with pause tokens.

Efficient Pretraining Length Scaling Think before you speak: Training language models with pause tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.763606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.634029Z digest=sha256:34af99ea34db4776194d42abcf908c58499f1fec8419c780a5eb44a88e4de0da

Observation 92849502-2711-40bf-ae22-62a82284aab2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Efficient Pretraining Length Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.638548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.638548Z digest=sha256:78cf8badf1052b295eacdb0a4fa61ed42de3132be1727ede737005984164e382

Observation 3f7478eb-756e-4cf2-aecd-d02af900b175 · outbound

This paper cites PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models.

Efficient Pretraining Length Scaling PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.642741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.642741Z digest=sha256:22aec66ae4cd6db23f0c345ae4b2673afdbd7943e7636aa50defface3181d780

Observation 80155363-9798-4441-97c7-6b2949a45db3 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Efficient Pretraining Length Scaling Training Large Language Models to Reason in a Continuous Latent Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.646890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.646890Z digest=sha256:cb6001e6a5c97e93c941fbff9162dec73bd11aeb3c558d6281a55051a5ea9ea5

Observation 7afb1830-a59a-4eaf-b75e-33d7f4ffe78f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient Pretraining Length Scaling Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.651372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.651372Z digest=sha256:64aa7f1627bed63bc429cd67770f95d7540899171232e02a24162ea6a46bf5d4

Observation 209e70b9-66b9-4053-bb77-0f80cd91d6c9 · outbound

This paper cites Scaling Laws for Transfer.

Efficient Pretraining Length Scaling Scaling Laws for Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.655454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.655454Z digest=sha256:77d7692a592971112b3e804ea6fa124e3a08a45d796ffda019a44cad3167d538

Observation f1b9f26b-e458-4b6f-9015-6ba6d282e343 · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

Efficient Pretraining Length Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.659632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.659632Z digest=sha256:d048285d77c87cd1cb14018c6866c4a7f7d742449646db8b9377d8ff231b68c5

Observation 163c4d11-ee79-453a-af08-6709b56dad43 · outbound

This paper cites Mistral 7B.

Efficient Pretraining Length Scaling Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.664264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.664264Z digest=sha256:45772367a3ddfd897c2142e289aa99e95637ba6051a05b5e0884d35caddd414a

Observation ca625383-aefd-4784-9ef3-633c5428580e · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024.

Efficient Pretraining Length Scaling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.669416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.669416Z digest=sha256:60f78aeec4e0f26e0a468c7be6ba5ef56a4ed8c9e2bda311b3d88a16cae9105b

Observation a06f4371-7320-412d-91cd-c08c3daec77f · outbound

This paper cites Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Efficient Pretraining Length Scaling Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.673545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.673545Z digest=sha256:36324de41cfe379c2f5d08047e7406a6653bdbe94e364978495e7ded61233a4c

Observation 497d90ed-2ef1-463c-818a-3a55a1e04edd · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Efficient Pretraining Length Scaling SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.677625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.677625Z digest=sha256:73d7e836c5ada1034e5f11d6b533c73e609037575a65159eed8e18afd390ceb5

Observation 7e232f65-c2fc-48fb-afb6-a44fe20f03a4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Efficient Pretraining Length Scaling Scaling Laws for Neural Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.681973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.681973Z digest=sha256:e7afb371e19b8911dfac19a56db4eb476e3a695fa0838ff1e3584fa9309502f6

Observation 1a5ddb20-51a2-40f5-a640-73dd082d3a0f · outbound

This paper cites Natural questions: a benchmark for question answering research.

Efficient Pretraining Length Scaling Natural questions: a benchmark for question answering research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.686173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.686173Z digest=sha256:d037bc15cb895a0c75e555bf831e503bfacac7473eba171870cd07688b8ee0a1

Observation 3437b068-4646-4092-950c-17ae22082461 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Efficient Pretraining Length Scaling Efficient memory management for large language model serving with pagedattention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.690438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.690438Z digest=sha256:24f07b1bc0101e4079c0172846c85ea07f2b0c2e8f0409b279b2cf7dcf8432e0

Observation a9cee53b-3d09-4738-9c19-20f8869de968 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

Efficient Pretraining Length Scaling MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.694729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.694729Z digest=sha256:f7290108a9b1fced14d7e475eae4a70d2af9d26fe60b133ab11eb9833025a6c3

Observation 31d8bb76-a3fe-4582-a07b-bc8dbe2f1d63 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.698916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.698916Z digest=sha256:324c893658596b849fc988daa5b2c6c062592888854ced75569f161de08dec75

Observation 689cae58-4f74-4dae-b1f9-6d8c08e0d3ea · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Efficient Pretraining Length Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.708789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.708789Z digest=sha256:a2e6de11df093e527b5726389ed15c3cc25d7bd026f390588a09ab8ff2895b3d

Observation e171f76d-316f-4c7e-ac7e-db93ce48ec98 · outbound

This paper cites DeepSeek-V3 Technical Report.

Efficient Pretraining Length Scaling DeepSeek-V3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.713440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.713440Z digest=sha256:aec814c8a5bcf61fd77d6f55b085222a31741c8b9364415ba76a64fac69fb09c

Observation b540b48a-e3b9-4079-bcf0-63075bde9b1f · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Efficient Pretraining Length Scaling Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.717780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.717780Z digest=sha256:df0e13fbb424d91649143e00c2a8846b6ce2ef872d82b87c56c11aba1e1a5bac

Observation 6970b460-588e-45a6-a150-d711cd57ba0a · outbound

This paper cites Cotformer: More tokens with attention make up for less depth.

Efficient Pretraining Length Scaling Cotformer: More tokens with attention make up for less depth

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.733667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.722872Z digest=sha256:a722232a04f2fb59f12f729e240b03360f34e82dfff5e54e9f4f646ffe191348

Observation 10601342-b229-4e83-bc93-b1f013364f00 · outbound

This paper cites 2 OLMo 2 Furious.

Efficient Pretraining Length Scaling 2 OLMo 2 Furious

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.726822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.726822Z digest=sha256:ed6209febf23f976185cfd8ecb1e6abbb65270b9ca5553c552b40392e8222d13

Observation cf7b5a91-4723-42d9-911b-62c308fbc876 · outbound

This paper cites Learning to reason with llms, 2024.

Efficient Pretraining Length Scaling Learning to reason with llms, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.730983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.730983Z digest=sha256:cf7e001be0482250fd6d332c4da9a9cf0b73daa5c168d03f90776ab0f7310074

Observation e15b3320-6f79-4d74-be8f-d37b62ddbe0a · outbound

This paper cites Learning to reason with llms, 2025.

Efficient Pretraining Length Scaling Learning to reason with llms, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.735013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.735013Z digest=sha256:b54f3c61fc76e4efbbad8c0511ec3d1d72456ddf9ed19396ba043a03508d9dc6

Observation 9e485649-1b75-42d9-b5bf-ab6eba06eb3f · outbound

This paper cites Vicky Zhao, Lili Qiu, and Dongmei Zhang.

Efficient Pretraining Length Scaling Vicky Zhao, Lili Qiu, and Dongmei Zhang

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-16T11:40:27.738918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.738918Z digest=sha256:c6a2329302426fd8bfc7e6f600d59d37d444c68484f646a962a0c52764df7cf1

Observation a340972b-df2c-4a84-a40c-c1ea645abdda · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Efficient Pretraining Length Scaling Gpqa: A graduate-level google-proof q&a benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.742943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.742943Z digest=sha256:21470c8185bdf25ce84fb9c469cd1c4e6a3a9d457fe4a769859149aa5380f07d

Observation e4498814-6522-4633-b142-e43347dd63db · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Efficient Pretraining Length Scaling SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.747393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.747393Z digest=sha256:7e3772c10daf2bb875d63f8baac0f1f614e7fc31ae510b957518b77fa5e60216

Observation 91e7e568-39a5-4459-84a7-f28fa8143952 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

Efficient Pretraining Length Scaling Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.751446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.751446Z digest=sha256:9c26764296e4f76f48304523e826b570c3b48eca4770429ed81279f7e37ea58a

Observation 942dcaf6-5bca-4dae-a22d-3cede4caf60c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Efficient Pretraining Length Scaling Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.755193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.755193Z digest=sha256:d6717dbbf69a791f18ea6c07e1a22f6c6b4f67266bf445c53e795488e165b3b9

Observation 5ae0e7c7-7d5d-4826-8039-6c42aebce566 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Efficient Pretraining Length Scaling FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.759136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.759136Z digest=sha256:1abcc424430f3094827e68a43dfdd03805867b8b7a36c6d1b2f92a4ac8c1a513

Observation a1ac3a2b-439b-4ba0-b529-fe90cf13ed15 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Efficient Pretraining Length Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.763275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.763275Z digest=sha256:d6be630fbb0154c9277b5679bf15d1040be9a8891a50785ea0599cf93f377b77

Observation cdf1eecf-8b62-4f5e-81d5-ed42dee80707 · outbound

This paper cites Sparsebert: Rethinking the importance analysis in self-attention.

Efficient Pretraining Length Scaling Sparsebert: Rethinking the importance analysis in self-attention

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.676571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.767431Z digest=sha256:21bd928f087ba2c32d15d197a8c333e5731219a89eead5e63d896e6812b4ee25

Observation ffa160b7-62b5-4451-9cee-ba8cc6e5f20b · outbound

This paper cites LLM Pretraining with Continuous Concepts.

Efficient Pretraining Length Scaling LLM Pretraining with Continuous Concepts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.771507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.771507Z digest=sha256:622a74386088287de8a06411b97966dd5d2be8ea0dbe5020b3ca5ec36c8179a5

Observation 9651b31e-bf8a-4581-8fe1-d93d9bb3bc07 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Efficient Pretraining Length Scaling CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.775667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.775667Z digest=sha256:1c4703d6a18dd0e659f9561cf487112d6045d5beecbfc4262795bbb1ba140576

Observation 62ff530d-9287-492f-8f52-d4e11f830858 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context llm inference.

Efficient Pretraining Length Scaling Quest: Query-aware sparsity for efficient long-context llm inference

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.663361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.780307Z digest=sha256:801c3d300423ff591b31299a713714c8e9280d53b84cd7ee4e7cee7039d9ef4a

Observation 9cf96c16-855e-49a5-a497-93f7bf9a3a39 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Efficient Pretraining Length Scaling Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.784477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.784477Z digest=sha256:87e4ff81cf3a192693dd8bc1eb97ed825de7e70f272548a652d01a72e8619d48

Observation f878e569-4991-4ae0-9492-22124a6c8e85 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Efficient Pretraining Length Scaling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.788641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.788641Z digest=sha256:6538c80bdd943b5ca72bc5843c9b3503e2ec14f0786fd728401f8e800527c03e

Observation d7008f2f-ee9d-4ac4-906b-d724213ddca8 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning.

Efficient Pretraining Length Scaling Spatten: Efficient sparse attention architecture with cascade token and head pruning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.792728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.792728Z digest=sha256:a098d2982c2353ae0e20c70f05e493e0a3c587975b2d85583f7a4ff7756e2d9f

Observation 087072a8-55d2-4a7f-9dfd-b557ca3f5767 · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents.

Efficient Pretraining Length Scaling Openhands: An open platform for ai software developers as generalist agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.796755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.796755Z digest=sha256:2eee7e471a0425ae64c44d5999193211f091af138fe38e9c7619243c8f0c5522

Observation 506aa59f-7d7a-4a29-8cb5-61cd03e2d6d7 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022.

Efficient Pretraining Length Scaling Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.800794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.800794Z digest=sha256:0cfa4165e3e5fe77c6865780bb697b86bfbab653f02dfaceb7d4d24666854ada

Observation b9b29f29-a0e1-4ad3-af00-201c0d2f4ad6 · outbound

This paper cites Efficient streaming language models with attention sinks.

Efficient Pretraining Length Scaling Efficient streaming language models with attention sinks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.633681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.804971Z digest=sha256:4a340971e641e59dabd622220dcdff8e150f0af66f80800cd0267fa6aee9c7dd

Observation 742911d0-84db-4836-a71b-037ca41da3bb · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Efficient Pretraining Length Scaling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.808920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.808920Z digest=sha256:51ccfb38649e25c8c2000180ac8e1eee27dfd1df5a8ad30e62a775838b488516

Observation 2cb9823b-4fc8-4f12-bf00-ca70efdd8d24 · outbound

This paper cites Big bird: Transformers for longer sequences.

Efficient Pretraining Length Scaling Big bird: Transformers for longer sequences

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.620016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.813063Z digest=sha256:f0da640c35964d95fe1e91efc2a38f508c2a6c4ba079b0c0a737af1ad1af1fa4

Observation fccd7f9e-4885-43b9-83c7-359895c4e452 · outbound

This paper cites Quiet-star: Language models can teach themselves to think before speaking.

Efficient Pretraining Length Scaling Quiet-star: Language models can teach themselves to think before speaking

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.606051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.816971Z digest=sha256:666906f791d2f7065578356189f7d453c59bda3d412216030709a58dd536a289

Observation d34e39bc-7abd-4d54-8354-81e1d36388c6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Efficient Pretraining Length Scaling HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.820980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.820980Z digest=sha256:4393034212de90e2479fe4968f19b71a5727cd8a3806641927dcd2653fd37558

Observation 41be79ca-507e-42dd-9121-b032ca2660bb · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Efficient Pretraining Length Scaling H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.592405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.825037Z digest=sha256:522c0fca3318e8273c0d7f01dcee7e435f3ec1e4cff756df71bd61bc671b47a6

Observation 100ce4b3-c7de-4b4b-953b-e3426844ab92 · outbound

This paper cites Accessed: 2024-9-29.

Efficient Pretraining Length Scaling Accessed: 2024-9-29

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.776898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:40:27.621049Z digest=sha256:99bd62fd0bf4afc2d6d26f5dd6b895b54d5481857bc3be31e9561cc867ac2822

Observation 2cbdfd7b-e0f9-45b7-95e4-990a0b58189b · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.704329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.704329Z digest=sha256:d3d8c144afe76207e3aedbd433d2a97fb34b99a868107ff15c162c625f6297b7

Pith citing papers

Observation 9e028773-dd76-486a-a7fa-06f751cc643f · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:43:11.770485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:42a70706670a836829f6c09065356c12bebddbfa0bed4aa2bbc6b38028c9bc7b

Observation b3c40fe4-0676-45ed-83ed-00fe7703e31f · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:30.334937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:30.334937Z digest=sha256:3257a639b5bd5c9261f0f120aafe0f292c70df60d5432b621f40020e0d7f922d

Observation 7003648e-fb23-457e-9306-e412ae4fdc60 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Efficient Pretraining Length Scaling

Reference 235

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:647222762681ac3adfcca8aa9d0e99d9126cb3e723ea55e2cfb7509e165b43c5

Observation f683faf2-3189-4eec-9029-366a612255d5 · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient Pretraining Length Scaling

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.285778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:092b33f6acbdcb43739bac0a0122078e6ed173b13ec7f2360c6d82db64cd4599