Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.825037Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2504.14992.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.825037Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:31:30.334937Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T07:43:11.766620Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2a495a46-b68e-4c72-b08b-8ed75d565a47 · outbound
Efficient Pretraining Length Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f726f6-5bd4-452b-b1ee-51c5db3e5e8e · outbound
Efficient Pretraining Length Scaling Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b54ffb-3f76-4735-b1f1-a3cc8cfe9928 · outbound
Efficient Pretraining Length Scaling Longformer: The Long-Document Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e29e45-f8e2-4088-83a9-6c5861deb028 · outbound
Efficient Pretraining Length Scaling Piqa: Reasoning about physical commonsense in natural language
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091c1abf-aee0-4067-8db9-777408e2aedd · outbound
Efficient Pretraining Length Scaling Striped Attention: Faster Ring Attention for Causal Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f1e1f3-92a8-43a5-b4ed-aa5dc82d590b · outbound
Efficient Pretraining Length Scaling Step-level Value Preference Optimization for Mathematical Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ecf90b-d3f6-4c55-ae68-2d61dd38abc3 · outbound
Efficient Pretraining Length Scaling Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98a2357-b202-48e5-ad30-4a0bf0bd91ea · outbound
Efficient Pretraining Length Scaling Generating Long Sequences with Sparse Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e293a4-de22-4678-9395-6af6dcc67457 · outbound
Efficient Pretraining Length Scaling Unified scaling laws for routed language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80048c91-2a95-4320-867c-e50c8cc87163 · outbound
Efficient Pretraining Length Scaling Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665c3ae4-19f4-4b9f-bc13-659ac7dbb42b · outbound
Efficient Pretraining Length Scaling Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8c855b-09c9-45d8-a140-9ace5b336d3f · outbound
Efficient Pretraining Length Scaling Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c858dc2e-bee8-45ca-b2d5-c2fbc0795e09 · outbound
Efficient Pretraining Length Scaling Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a35e5cf4-34aa-4230-be4f-26896c2c47d0 · outbound
Efficient Pretraining Length Scaling Flash-decoding for long-context inference, October
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0fcddfd9-b261-4dee-aa43-098683fe5ded · outbound
Efficient Pretraining Length Scaling LongNet: Scaling Transformers to 1,000,000,000 Tokens
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fe73a9-05d6-4945-bf02-11a2bcb339ef · outbound
Efficient Pretraining Length Scaling Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26e4082-807b-40ae-a037-b99d19b14464 · outbound
Efficient Pretraining Length Scaling Think before you speak: Training language models with pause tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92849502-2711-40bf-ae22-62a82284aab2 · outbound
Efficient Pretraining Length Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7478eb-756e-4cf2-aecd-d02af900b175 · outbound
Efficient Pretraining Length Scaling PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80155363-9798-4441-97c7-6b2949a45db3 · outbound
Efficient Pretraining Length Scaling Training Large Language Models to Reason in a Continuous Latent Space
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afb1830-a59a-4eaf-b75e-33d7f4ffe78f · outbound
Efficient Pretraining Length Scaling Measuring Massive Multitask Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209e70b9-66b9-4053-bb77-0f80cd91d6c9 · outbound
Efficient Pretraining Length Scaling Scaling Laws for Transfer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b9f26b-e458-4b6f-9015-6ba6d282e343 · outbound
Efficient Pretraining Length Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163c4d11-ee79-453a-af08-6709b56dad43 · outbound
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca625383-aefd-4784-9ef3-633c5428580e · outbound
Efficient Pretraining Length Scaling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06f4371-7320-412d-91cd-c08c3daec77f · outbound
Efficient Pretraining Length Scaling Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497d90ed-2ef1-463c-818a-3a55a1e04edd · outbound
Efficient Pretraining Length Scaling SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e232f65-c2fc-48fb-afb6-a44fe20f03a4 · outbound
Efficient Pretraining Length Scaling Scaling Laws for Neural Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5ddb20-51a2-40f5-a640-73dd082d3a0f · outbound
Efficient Pretraining Length Scaling Natural questions: a benchmark for question answering research
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3437b068-4646-4092-950c-17ae22082461 · outbound
Efficient Pretraining Length Scaling Efficient memory management for large language model serving with pagedattention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9cee53b-3d09-4738-9c19-20f8869de968 · outbound
Efficient Pretraining Length Scaling MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d8bb76-a3fe-4582-a07b-bc8dbe2f1d63 · outbound
Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689cae58-4f74-4dae-b1f9-6d8c08e0d3ea · outbound
Efficient Pretraining Length Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e171f76d-316f-4c7e-ac7e-db93ce48ec98 · outbound
Efficient Pretraining Length Scaling DeepSeek-V3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b540b48a-e3b9-4079-bcf0-63075bde9b1f · outbound
Efficient Pretraining Length Scaling Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6970b460-588e-45a6-a150-d711cd57ba0a · outbound
Efficient Pretraining Length Scaling Cotformer: More tokens with attention make up for less depth
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 10601342-b229-4e83-bc93-b1f013364f00 · outbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7b5a91-4723-42d9-911b-62c308fbc876 · outbound
Efficient Pretraining Length Scaling Learning to reason with llms, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e15b3320-6f79-4d74-be8f-d37b62ddbe0a · outbound
Efficient Pretraining Length Scaling Learning to reason with llms, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e485649-1b75-42d9-b5bf-ab6eba06eb3f · outbound
Efficient Pretraining Length Scaling Vicky Zhao, Lili Qiu, and Dongmei Zhang
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a340972b-df2c-4a84-a40c-c1ea645abdda · outbound
Efficient Pretraining Length Scaling Gpqa: A graduate-level google-proof q&a benchmark
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4498814-6522-4633-b142-e43347dd63db · outbound
Efficient Pretraining Length Scaling SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e7e568-39a5-4459-84a7-f28fa8143952 · outbound
Efficient Pretraining Length Scaling Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942dcaf6-5bca-4dae-a22d-3cede4caf60c · outbound
Efficient Pretraining Length Scaling Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae0e7c7-7d5d-4826-8039-6c42aebce566 · outbound
Efficient Pretraining Length Scaling FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ac3a2b-439b-4ba0-b529-fe90cf13ed15 · outbound
Efficient Pretraining Length Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf1eecf-8b62-4f5e-81d5-ed42dee80707 · outbound
Efficient Pretraining Length Scaling Sparsebert: Rethinking the importance analysis in self-attention
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ffa160b7-62b5-4451-9cee-ba8cc6e5f20b · outbound
Efficient Pretraining Length Scaling LLM Pretraining with Continuous Concepts
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9651b31e-bf8a-4581-8fe1-d93d9bb3bc07 · outbound
Efficient Pretraining Length Scaling CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ff530d-9287-492f-8f52-d4e11f830858 · outbound
Efficient Pretraining Length Scaling Quest: Query-aware sparsity for efficient long-context llm inference
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9cf96c16-855e-49a5-a497-93f7bf9a3a39 · outbound
Efficient Pretraining Length Scaling Gemini: A Family of Highly Capable Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f878e569-4991-4ae0-9492-22124a6c8e85 · outbound
Efficient Pretraining Length Scaling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7008f2f-ee9d-4ac4-906b-d724213ddca8 · outbound
Efficient Pretraining Length Scaling Spatten: Efficient sparse attention architecture with cascade token and head pruning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087072a8-55d2-4a7f-9dfd-b557ca3f5767 · outbound
Efficient Pretraining Length Scaling Openhands: An open platform for ai software developers as generalist agents
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506aa59f-7d7a-4a29-8cb5-61cd03e2d6d7 · outbound
Efficient Pretraining Length Scaling Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b29f29-a0e1-4ad3-af00-201c0d2f4ad6 · outbound
Efficient Pretraining Length Scaling Efficient streaming language models with attention sinks
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 742911d0-84db-4836-a71b-037ca41da3bb · outbound
Efficient Pretraining Length Scaling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb9823b-4fc8-4f12-bf00-ca70efdd8d24 · outbound
Efficient Pretraining Length Scaling Big bird: Transformers for longer sequences
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fccd7f9e-4885-43b9-83c7-359895c4e452 · outbound
Efficient Pretraining Length Scaling Quiet-star: Language models can teach themselves to think before speaking
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d34e39bc-7abd-4d54-8354-81e1d36388c6 · outbound
Efficient Pretraining Length Scaling HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41be79ca-507e-42dd-9121-b032ca2660bb · outbound
Efficient Pretraining Length Scaling H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 100ce4b3-c7de-4b4b-953b-e3426844ab92 · outbound
Efficient Pretraining Length Scaling Accessed: 2024-9-29
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2cbdfd7b-e0f9-45b7-95e4-990a0b58189b · outbound
Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e028773-dd76-486a-a7fa-06f751cc643f · inbound
Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b3c40fe4-0676-45ed-83ed-00fe7703e31f · inbound
Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7003648e-fb23-457e-9306-e412ae4fdc60 · inbound
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Efficient Pretraining Length Scaling
Reference 235
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f683faf2-3189-4eec-9029-366a612255d5 · inbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient Pretraining Length Scaling
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.