Pith. sign in

Paper Citation Record · LEDGER

Scaling Inference-Efficient Language Models

As of 13 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2501.18107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18107 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:43:29.577850Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:08:06.751426Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:15:31.617791Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08438660-135f-4f83-8318-11db118c05cf · outbound

This paper cites Results over vLLM In this section, we first evaluate the inference efficiency of open-source large language models over vLLM using NVIDIA Tesla A100 Ampere 40 GB GPU.

Scaling Inference-Efficient Language Models Results over vLLM In this section, we first evaluate the inference efficiency of open-source large language models over vLLM using NVIDIA Tesla A100 Ampere 40 GB GPU

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.440594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.561143Z digest=sha256:7acc85083ebdf5d2c9ef23bbc81ef181199d1a280f137d3695e27f64177de469

Observation 2b1402f3-23ec-4a99-b4c5-b2d7da5810f7 · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

Scaling Inference-Efficient Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.592940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.592940Z digest=sha256:8de60a3bf4cd75727b8a82043c93123d5d7dcbd243a269e2893839e0c9526935

Observation e99b4b80-e0aa-44e4-bb15-535bb83e9156 · outbound

This paper cites dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads.

Scaling Inference-Efficient Language Models dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.521770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.422891Z digest=sha256:9bcfb0ed2462a1fd8814663ccd8a65bebab3be8fd229482f7b905a23535fd190

Observation f8420ad8-50cf-478b-8687-4f560a193fe0 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Scaling Inference-Efficient Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.620437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.620437Z digest=sha256:aff8125bbf6e99119b85869e8f3bdfa0aebedaad78884d554736c20ad64617db

Observation a6c22c8c-5962-4fbf-8814-178c017b4dad · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Scaling Inference-Efficient Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.624251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.624251Z digest=sha256:774f39659989dcb386e020e6cd2b6e28b37a58a505763e4f2b4984aa5f0d274f

Observation 41086873-bfa6-4052-9826-7283fb07b62a · outbound

This paper cites The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024).

Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.510314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.459723Z digest=sha256:36f72ce5a32c750077ae1fcddd9644ba723fe93a628f13d2781040ff9006c57c

Observation 9acfc0b0-95f6-439f-97a1-f5df976f7a35 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Scaling Inference-Efficient Language Models Language models scale reliably with over-training and on downstream tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.854869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.854869Z digest=sha256:3da203d74c95115558990d03ff6cd95fab6dff8d4aece9e5914fef6918184698

Observation 8a484819-c134-4763-9cce-a00b0300d799 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

Scaling Inference-Efficient Language Models SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.887469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.887469Z digest=sha256:30a82e8a8a2ae8ebcda146f5f24c78c21ba35d555a7ae158984c8853dd56ea19

Observation 034e27fd-824d-48e9-bf03-ea75755be632 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Scaling Inference-Efficient Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.905055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.905055Z digest=sha256:249509deb5b4ed074e54a16ee9e86a4e9175a4d9147157717a6b161e7e7ed16c

Observation cb9179df-6538-4e28-954a-eebf8fb915fa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Inference-Efficient Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.919255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.919255Z digest=sha256:5289abb38319e8700fa6bf321d708ae2842819052ae6fafbc8029e50e3d2b8a2

Observation ad7d11a5-4b1c-45e8-946d-01b54697d36b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Inference-Efficient Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.925640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.925640Z digest=sha256:30cb93d0a8565c3f03413cde2bda34586f5f78a0bf4687c6d65671d32ef1aaa0

Observation d53c4378-4f48-474a-af51-15bfe863e5a9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Inference-Efficient Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.929669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.929669Z digest=sha256:db124dcd7c29a13dab56136420ee2816af58ee481a2481aa442c08ccde129f65

Observation af18da9b-6d42-4283-9145-de574b0b3f52 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Inference-Efficient Language Models Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.933586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.933586Z digest=sha256:7ad81120327eeda0716ee44fcf6570fae7a5c6f5c3b2cce62bc1c22602879084

Observation 54d7b0e2-73b2-4a48-a3b9-f20cc31d373c · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Scaling Inference-Efficient Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.937752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.937752Z digest=sha256:278816b3604566eef4cbdc6edb2db330fc3ece57182a6bab720d9849aa07449a

Observation 5eb28230-aa94-400e-806e-1f4f1ac094ca · outbound

This paper cites OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization.

Scaling Inference-Efficient Language Models OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.941174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.941174Z digest=sha256:e2ed5e9c782de10bababbb935d579d282eaf4168d6526aee4b6676c82eaf4db7

Observation d2582b16-804f-4781-8624-a8375140aac3 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Scaling Inference-Efficient Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.945003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.945003Z digest=sha256:6be690d14ce4388c831e34cf4f420d56a69505489ffcc9edc43af2ae7c406f2f

Observation b79f3d96-6d66-44fc-abbb-039d5df1ac2a · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Inference-Efficient Language Models Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.949049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.949049Z digest=sha256:d673483920563c40adae90d1e02f9b39a38e0bc4b4ad612040486cc766af8fd8

Observation 743d0cea-2e11-4eb2-99ef-552c3de26c33 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Scaling Inference-Efficient Language Models Scaling Laws for Fine-Grained Mixture of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.952649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.952649Z digest=sha256:2af35899be390fff9a6837fee1440b9794e8dc84150e7d2e4711b5d01c52700d

Observation 666f646f-b6b1-47c6-95dd-a7f66d8fcb67 · outbound

This paper cites Scaling Laws for Precision.

Scaling Inference-Efficient Language Models Scaling Laws for Precision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.956199Z digest=sha256:3db25e4d45dbd6e96aae32bce6e76695527ccafcde2a23590eca20ee30d0cb1e

Observation 779028af-5f12-49e8-b0c4-84a3dd7a88cc · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Scaling Inference-Efficient Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.960173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.960173Z digest=sha256:e42353651f9c90dbd6ecb31a091acb857c6c1c415e44f9e0e229e86e225a5ba3

Observation 9dffec26-ad04-42ff-8152-7cac27223013 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Scaling Inference-Efficient Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.964120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.964120Z digest=sha256:0b2216fdaab5015cd254ba8975719b0014f116341eab7980c24a692e2ba56d3e

Observation 2d390137-7835-4a37-bb3a-ab1eb0fed272 · outbound

This paper cites MLC-LLM, 2023-2025.

Scaling Inference-Efficient Language Models MLC-LLM, 2023-2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.544383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:28.967586Z digest=sha256:d572fed4420b3f72854ae7fb43708847a96b29caefccf7465820fee5e65a406e

Observation c39a3c32-edfa-4498-a651-cb7a50cd9876 · outbound

This paper cites TensorFlow-Serving: Flexible, High-Performance ML Serving.

Scaling Inference-Efficient Language Models TensorFlow-Serving: Flexible, High-Performance ML Serving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.971214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.971214Z digest=sha256:ca691b19d65ab74590f6bdc257fb5e0887ba8985e6b103f099abf8737c4db9ff

Observation d03d123a-ee22-4941-be21-0958ae264f8a · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Inference-Efficient Language Models Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.139416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.139416Z digest=sha256:c14ae3a581fdec85a7dd2af8d6c3d29c92f5a059d614570751c075cfd2548f2a

Observation 2b19d0a3-c7a2-490a-9d9e-b3bd645aac06 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Scaling Inference-Efficient Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.171747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.171747Z digest=sha256:5b0dcd67176ea84bb3553972449f83c23738778f79fe5918555c31f9a53d6e45

Observation 8097040e-2e6c-40b2-a1bd-28c807b7bda2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Scaling Inference-Efficient Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.181113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.181113Z digest=sha256:84fb7e4b460ef18f142aaaa30c880cc81a98fc171b0329838577b53855d36f55

Observation 65e47aa4-c5d0-4a68-944d-b62ee373bee7 · outbound

This paper cites Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies.

Scaling Inference-Efficient Language Models Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.213990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.213990Z digest=sha256:8ea237929283797ac1f097b952f9506c9f424b4c1df16b7ff4797c8ba8c43a73

Observation dc59e599-d3da-4b64-9257-7944773488cd · outbound

This paper cites Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers.

Scaling Inference-Efficient Language Models Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.237192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.237192Z digest=sha256:50b39fc67a090424a48fb087248819a8a2779a9e234d339b9a1a448ce7d34c14

Observation 8bfdace5-9950-45a0-85f6-edef190d15f8 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Inference-Efficient Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.244586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.244586Z digest=sha256:350b037b1ca53a96a55d57a12e9aa7bcaf3761a938906ca4151f4c5cea02ff82

Observation 1a995995-f6c7-451f-a4e0-6671d0b5e10b · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Scaling Inference-Efficient Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.248752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.248752Z digest=sha256:43947bc37528cf28a03770d4a322ad39ad11476653c0e512dd2c4c63b2e91faf

Observation 320db702-f16d-48ef-9b86-197de543fb9e · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Scaling Inference-Efficient Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.252725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.252725Z digest=sha256:28fa79bbaabe5fcb1eacb267178b0d55fbf2f9441e64162029e27c022b676372

Observation e8361014-a9c4-4b52-ae56-22eb6041ae8c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Scaling Inference-Efficient Language Models DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.260648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.260648Z digest=sha256:146a41dd1d00d993a6de11cd43cad2a70ae5736641da8023df675928165ba291

Observation 79f4d5c7-acd7-4733-ac4d-20ae0c18214a · outbound

This paper cites Decoding Speculative Decoding.

Scaling Inference-Efficient Language Models Decoding Speculative Decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.264303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.264303Z digest=sha256:c609113b597108dab00142c097c522b2c76037ae6287d01f9a0873fce1a70330

Observation 130af53b-c465-461f-ad57-d562fce51fa9 · outbound

This paper cites Qwen2.5 Technical Report.

Scaling Inference-Efficient Language Models Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.268157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.268157Z digest=sha256:04e31839f31ea12f7586e5e4484bb4711b52053f0336e361fb450897a7cf6695

Observation 6edbb00f-43ea-488c-9d79-e57a745fb3d2 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Scaling Inference-Efficient Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.271945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.271945Z digest=sha256:609b3ee8220e8d9093e2ec57d43c814aeef69337be6f1f32ed71740e4d18c9a0

Observation 71770aef-bfb3-4c91-8934-b4b558f885c0 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

Scaling Inference-Efficient Language Models FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.275583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.275583Z digest=sha256:74ebe1721488caeac3f1add24f5650b63e6a2c53dceadf9075eaaf5cec3c1c41

Observation 7e502ea8-cd4f-44bb-a91c-8aacd2cf5bc5 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Scaling Inference-Efficient Language Models Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.279768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.279768Z digest=sha256:2580ab07dc7490de99e37f9a0f8595547845c8371da74552a93a9186e9a7ecc8

Observation 00bc4bce-51e1-4033-a47d-bff38ca9befd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Scaling Inference-Efficient Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.317998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.317998Z digest=sha256:d15cad5b87b72a80bfc3a1d0841f3053adf157367ab357dbb67ae5a6984675bd

Observation f36b0fb5-36bf-46ce-bafd-12008f918a7d · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Scaling Inference-Efficient Language Models OPT: Open Pre-trained Transformer Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.358954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.358954Z digest=sha256:d929f14e66fe0617e1cd8a018f157bb9e30cd101b62a4badda69e57f54c27361

Observation 9809848f-786e-40cc-b902-b7d36b2a17df · outbound

This paper cites (Center) We indicate the relationship between inference latency and hidden size with the number of layers fixed.

Scaling Inference-Efficient Language Models (Center) We indicate the relationship between inference latency and hidden size with the number of layers fixed

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.471291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.557585Z digest=sha256:67f099c072da42590e72fdf15055b0408d0a784965283286f9036cfec63804a0

Observation 92eb9d33-d361-4bba-9201-2c641ead1be5 · outbound

This paper cites The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024).

Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.400895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.564505Z digest=sha256:bced8f8c8b586aadd5060a88a320f6f5a48a7081ab498bd901f806a00949af77

Observation c0cfc69b-3642-40db-ac9d-47249834c10f · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.324478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.567890Z digest=sha256:32b235c6a078dd4c0c8b165d1bee2947574287167f3eb40c2f04dbfa0c3df2cf

Observation 8b000cb5-25d9-4ce2-bc4c-40069c0d50c6 · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.178497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.571468Z digest=sha256:7001299734c8c0386635a147e3a19bd2ba071f548f646927570f6b9fdd0d6dfc

Observation a6da3976-01ab-456d-9366-3039c7b87aaa · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.134727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.574447Z digest=sha256:ae3d6875333bd3788fb9526118bb524cdb3202fce89c5dc8509c7d0da1196e13

Observation 459ddbfe-1a6c-4f01-bc00-a4463a9c4683 · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.123552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.577850Z digest=sha256:d99bed225bb597978a003f99cda1af3d25640466348abdb89b03786928e1fc0e

Observation e805e0bf-72d8-4660-b255-d504e1315cfe · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.496911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.537141Z digest=sha256:fd06c97a5d8ee21079140f0ca59e12c8c2d87eb442bad41605b29993b422891d

Observation 6266e32d-d3a6-4a57-b7b7-76549fca7d7b · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Scaling Inference-Efficient Language Models Efficient Streaming Language Models with Attention Sinks

Reference 1921

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.256420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.256420Z digest=sha256:f715a44177e6a72d8da9d39fb5f401fe8e1259269b84f7600f4b11255c07b401

Observation ef6e98a8-048f-41af-abf5-fa1949c0175f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Scaling Inference-Efficient Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 1961

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.198924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.198924Z digest=sha256:7af98f6a4c746678e511a79c4f81279423cfcbcfcf662b2f3dadeed0032e2922

Observation 72478f8f-797d-4a77-8054-4166b6a683fe · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Scaling Inference-Efficient Language Models Observational Scaling Laws and the Predictability of Language Model Performance

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.116548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.116548Z digest=sha256:23c3f11b6cf6b478f155338017a7e27ebba80c916c63283887b53b773af381e6

Observation 5f881308-29a5-4563-ba0e-b23af1b6e46f · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Scaling Inference-Efficient Language Models Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.080074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.080074Z digest=sha256:b08cf21ddf0c959fc9a05e6d359ab4272df0342defb0aa8ddf76720e720338db

Observation f4d96b4e-56c0-4bf1-98d8-0e3684f6fc53 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Scaling Inference-Efficient Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.016527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.016527Z digest=sha256:17c8bdfcfee59a8fdda50b754261fffb505f1cc3e5b0c9302bdcd7af62aaed2c

Observation c9fc260f-b242-46db-99aa-ed15d558a6b6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Inference-Efficient Language Models Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.768781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.768781Z digest=sha256:66baa127bd31a0590354340f6d10c0554141d8fc01feafe98dfe12758b659bee

Observation e3130bf6-75bc-4e61-824d-6417abde6ec9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Inference-Efficient Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.672091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.672091Z digest=sha256:bd0c1b356ae2247ed5fa312ebe1727664fafaf11a1cc4e66ecde69bc45ff6400

Observation 9fce7d57-3670-412a-b473-88a41959fe5e · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Scaling Inference-Efficient Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.612128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.612128Z digest=sha256:4f75a85292a448cce60aa1be58e2b84ff993c6bb30ca1ee40bdd74f40dc7c457

Observation f9b70a2b-b9ac-4d21-ad99-c0df5c2c6344 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Inference-Efficient Language Models The Llama 3 Herd of Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.819084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.819084Z digest=sha256:bd0d4694bb585674771a195dac2288c6e3d1145209e7261a911b8026ef958edd

Observation 764af2e5-a84a-46ec-ab84-f40ab30dc6e6 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Scaling Inference-Efficient Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.616404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.616404Z digest=sha256:4ade2b5c729f47e622996c02bcd0c4bde6c23442694e59aafe541799209023da

Observation 2a3c5454-eb55-4b54-8778-c051d3ff73bf · outbound

This paper cites Qwen Technical Report.

Scaling Inference-Efficient Language Models Qwen Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.607548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.607548Z digest=sha256:cf3266ca25b6e71a1bf9e7ef977c0945f200d3423833d173fe6d7f6b5d148073

Observation 615efc87-a32f-4aea-8c20-8bfd1f41ff8c · outbound

This paper cites Phi-4 Technical Report.

Scaling Inference-Efficient Language Models Phi-4 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.585716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.585716Z digest=sha256:6f30ef97de16f871e30a30035d16bd4af6fa60c83956da9ce1708c9811c07978

Observation 9965cdfd-08b6-47ba-9256-79a8b6f2a577 · outbound

This paper cites CHAI: Clustered Head Attention for Efficient LLM Inference.

Scaling Inference-Efficient Language Models CHAI: Clustered Head Attention for Efficient LLM Inference

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:43:30.086880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:28.601962Z digest=sha256:091774dddf5bb13264f39c6aebe2d34d404b444d750129f593145bfe8ceb2e7b

Observation 8e4fcd9c-f716-4138-b306-c9f8e6b8b30e · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 2048

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.533224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T00:43:29.385782Z digest=sha256:41d96ad43b8a189a351b1fbf21fc73568980747a7f5799d2769011d4b80b7de2

Pith citing papers

Observation 49760f12-0291-4357-a221-0caf1b0379f9 · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Scaling Inference-Efficient Language Models

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:40.215851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:40.215851Z digest=sha256:cc0cdf2ae35dbf5f0a32e6c71f7b53806bb79f176387ece8c0117e72d1c800b0

Observation b109582b-e94e-4858-853b-129ad6b632fc · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Inference-Efficient Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.965606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:828fa0f8f41b6861d3f2b0bcabc38623624a014cdc048d1ea5e11651d4bb324b

Observation ad67b25a-8e8c-41f2-9975-9de5ae1c9a41 · inbound

Comprehensive AI governance requires addressing non-model gains cites this paper.

Comprehensive AI governance requires addressing non-model gains Scaling Inference-Efficient Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:31.619579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T08:11:23.860723Z digest=sha256:7077a917c24d23f9646047a0f5e2e5e2f587f12ee1aafa1d9934f944c71d4950

Observation 58850515-0fa4-454f-b1c8-916159b0c19b · inbound

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts cites this paper.

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts Scaling Inference-Efficient Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:08:06.751426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:08:06.751426Z digest=sha256:e3b97da161d0715f978fa383a1b0dfa11bd1d571fae7381bbf78aa99bd4d1728