Pith. sign in

Paper Citation Record · LEDGER

Scaling Inference-Efficient Language Models

As of 10 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2501.18107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18107 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:43:29.577850Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:40.215851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:15:31.617791Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08438660-135f-4f83-8318-11db118c05cf · outbound

This paper cites Results over vLLM In this section, we first evaluate the inference efficiency of open-source large language models over vLLM using NVIDIA Tesla A100 Ampere 40 GB GPU.

Scaling Inference-Efficient Language Models Results over vLLM In this section, we first evaluate the inference efficiency of open-source large language models over vLLM using NVIDIA Tesla A100 Ampere 40 GB GPU

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.440594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.561143Z digest=sha256:845ceee382ee516efc51d2b23a0ae5e7dd21df8c0c8499f38501eb80b83a7786

Observation 2b1402f3-23ec-4a99-b4c5-b2d7da5810f7 · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

Scaling Inference-Efficient Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.592940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.592940Z digest=sha256:c2e266e16cf92e4b49293f77646d223c9bd1eedce73e2c6f0533fece5d746b43

Observation e99b4b80-e0aa-44e4-bb15-535bb83e9156 · outbound

This paper cites dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads.

Scaling Inference-Efficient Language Models dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.521770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.422891Z digest=sha256:703b328c323f6cca3cedb7e4da0d0a3912bfa3623f45d57e4a2b7a4b6ab50516

Observation f8420ad8-50cf-478b-8687-4f560a193fe0 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Scaling Inference-Efficient Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.620437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.620437Z digest=sha256:e085e135a80b351c9b6b31bb8f52c7590f0517e69560a87b5de49974c039076c

Observation a6c22c8c-5962-4fbf-8814-178c017b4dad · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Scaling Inference-Efficient Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.624251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.624251Z digest=sha256:7a31e67390f199414985054f3c2b55a867f7160bb931a7e719fd10ab97e7fbac

Observation 41086873-bfa6-4052-9826-7283fb07b62a · outbound

This paper cites The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024).

Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.510314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.459723Z digest=sha256:0b4b2d83547ca4befcb441a1ec18efdf811bae9029154d4ce0b6e7f4d905070e

Observation 9acfc0b0-95f6-439f-97a1-f5df976f7a35 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Scaling Inference-Efficient Language Models Language models scale reliably with over-training and on downstream tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.854869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.854869Z digest=sha256:1ef93e47b0507a4dae2c9b006d002fef26d49758db11151f4689f23cd9c67349

Observation 8a484819-c134-4763-9cce-a00b0300d799 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

Scaling Inference-Efficient Language Models SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.887469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.887469Z digest=sha256:8d7360cd80710d64797b5420ead8a6068bad4820a7e3132fe840c1c2b8a5a062

Observation 034e27fd-824d-48e9-bf03-ea75755be632 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Scaling Inference-Efficient Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.905055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.905055Z digest=sha256:4ece56f2ee29f4348534cdf4176c7a486588b9b1de9d776589793c35fc982b95

Observation cb9179df-6538-4e28-954a-eebf8fb915fa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Inference-Efficient Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.919255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.919255Z digest=sha256:32073a9c15f99475af093c8e395645c9115eff994f7958a8c51b6343cb973397

Observation ad7d11a5-4b1c-45e8-946d-01b54697d36b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Inference-Efficient Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.925640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.925640Z digest=sha256:7004055b3ffcb9b696d87cf56ed5b6245fc225352240a4725c8f8dbe8933a3c8

Observation d53c4378-4f48-474a-af51-15bfe863e5a9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Inference-Efficient Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.929669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.929669Z digest=sha256:0acf91ecb209d6d808542d58175577f2bf61bd0444aa56314238cb4d4b19c8db

Observation af18da9b-6d42-4283-9145-de574b0b3f52 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Inference-Efficient Language Models Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.933586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.933586Z digest=sha256:5f08f35c9dadef67f36ab2ae83f53f2cb612d07bfc05ddb50362da88b43aa126

Observation 54d7b0e2-73b2-4a48-a3b9-f20cc31d373c · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Scaling Inference-Efficient Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.937752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.937752Z digest=sha256:04c6d35d842dac82a44f1d96e100fc28b68c5559b009e816555e6a682f197a67

Observation 5eb28230-aa94-400e-806e-1f4f1ac094ca · outbound

This paper cites OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization.

Scaling Inference-Efficient Language Models OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.941174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.941174Z digest=sha256:6c3540ce1554db904d4e6af2ed1da3608473072c679f3da90d0d0b053216f04c

Observation d2582b16-804f-4781-8624-a8375140aac3 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Scaling Inference-Efficient Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.945003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.945003Z digest=sha256:9d6946f3ab3776fe7b9c09a039b1177ea03499f283ea6b82f3b72561477ac27f

Observation b79f3d96-6d66-44fc-abbb-039d5df1ac2a · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Inference-Efficient Language Models Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.949049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.949049Z digest=sha256:3525a915fd21afb607ebb98091ba6945a91d713303ece7f425486a188ba511aa

Observation 743d0cea-2e11-4eb2-99ef-552c3de26c33 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Scaling Inference-Efficient Language Models Scaling Laws for Fine-Grained Mixture of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.952649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.952649Z digest=sha256:9639a207454c86fbd24e37893f5479fa0cf21d236a5cd8f84cb4ed6ae3833209

Observation 666f646f-b6b1-47c6-95dd-a7f66d8fcb67 · outbound

This paper cites Scaling Laws for Precision.

Scaling Inference-Efficient Language Models Scaling Laws for Precision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.956199Z digest=sha256:1b45630637bb49638dc247e1f6790aa9b68eeeab002e55c36e6b0dde5b1fa0d6

Observation 779028af-5f12-49e8-b0c4-84a3dd7a88cc · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Scaling Inference-Efficient Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.960173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.960173Z digest=sha256:c181db746f7e6e8de942ee37e828aac92ec6804cfa85ed2d2307df8831379d6d

Observation 9dffec26-ad04-42ff-8152-7cac27223013 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Scaling Inference-Efficient Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.964120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.964120Z digest=sha256:55a8772c2042f67113ca2351429559e67102b13dba8e0640dd7b22d31fc1a6a0

Observation 2d390137-7835-4a37-bb3a-ab1eb0fed272 · outbound

This paper cites MLC-LLM, 2023-2025.

Scaling Inference-Efficient Language Models MLC-LLM, 2023-2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.544383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:28.967586Z digest=sha256:e98b5c0e1b4aaf77838732db61343520e92880aabaa14f2db73706766b015ac3

Observation c39a3c32-edfa-4498-a651-cb7a50cd9876 · outbound

This paper cites TensorFlow-Serving: Flexible, High-Performance ML Serving.

Scaling Inference-Efficient Language Models TensorFlow-Serving: Flexible, High-Performance ML Serving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.971214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.971214Z digest=sha256:d602a5a4666322ee729f42674b32090c39e39d8b7ca6613e6b4ef75dc3c2a38c

Observation d03d123a-ee22-4941-be21-0958ae264f8a · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Inference-Efficient Language Models Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.139416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.139416Z digest=sha256:ef2a47ae2fafaac0ac568ce256dea1079a411d2dea818bb6bc23911437bf716c

Observation 2b19d0a3-c7a2-490a-9d9e-b3bd645aac06 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Scaling Inference-Efficient Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.171747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.171747Z digest=sha256:1d7226037a27bad00a669d1331cc9caf951ceb436564d94806b7249e8ecb6f33

Observation 8097040e-2e6c-40b2-a1bd-28c807b7bda2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Scaling Inference-Efficient Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.181113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.181113Z digest=sha256:ae454df4d930c1074c1c4b9d6654e613e5dc9db8a09610713d14411185633c12

Observation 65e47aa4-c5d0-4a68-944d-b62ee373bee7 · outbound

This paper cites Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies.

Scaling Inference-Efficient Language Models Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.213990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.213990Z digest=sha256:dbfb7b17435659a8ae6b16bb2727ad21bfa6aa053191833e307db7f2514d5dd8

Observation dc59e599-d3da-4b64-9257-7944773488cd · outbound

This paper cites Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers.

Scaling Inference-Efficient Language Models Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.237192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.237192Z digest=sha256:c227721f3c8b1ee88161cbf864d4dad479278ad9ca0184c0d115081f4aac6138

Observation 8bfdace5-9950-45a0-85f6-edef190d15f8 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Inference-Efficient Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.244586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.244586Z digest=sha256:022ee92f686556d43639b8b37e81d3325fefbd9a1c3a3ea562c8e55da6793535

Observation 1a995995-f6c7-451f-a4e0-6671d0b5e10b · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Scaling Inference-Efficient Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.248752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.248752Z digest=sha256:baa4fd179401db5c18e9d5a29e84e0bfa2d13dcb240c86814d14a21f89452419

Observation 320db702-f16d-48ef-9b86-197de543fb9e · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Scaling Inference-Efficient Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.252725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.252725Z digest=sha256:26b697ab614f40f3501a7f1dbbbbe2e2aa0338174d16a2f98dc2f6f6402d2268

Observation e8361014-a9c4-4b52-ae56-22eb6041ae8c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Scaling Inference-Efficient Language Models DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.260648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.260648Z digest=sha256:c4048f2337f36445a8057b5b2fda66f518e7e8ea274c9057d845176274311c9a

Observation 79f4d5c7-acd7-4733-ac4d-20ae0c18214a · outbound

This paper cites Decoding Speculative Decoding.

Scaling Inference-Efficient Language Models Decoding Speculative Decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.264303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.264303Z digest=sha256:8197e9d284c7b86e2a87fdce3b0a1b0f262a9564f5f6f4525df678e5e68dce58

Observation 130af53b-c465-461f-ad57-d562fce51fa9 · outbound

This paper cites Qwen2.5 Technical Report.

Scaling Inference-Efficient Language Models Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.268157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.268157Z digest=sha256:1d4c4924e7d055027c99ba13514279481e53fcd6eb25d34e6b026be629b24c10

Observation 6edbb00f-43ea-488c-9d79-e57a745fb3d2 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Scaling Inference-Efficient Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.271945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.271945Z digest=sha256:a9a051ce24d50850da5e30dca98fc2c41a2d4ae6808f7eeb386e1da18025e312

Observation 71770aef-bfb3-4c91-8934-b4b558f885c0 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

Scaling Inference-Efficient Language Models FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.275583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.275583Z digest=sha256:fc7b66d440e9388bac3620b8163bd6981e65b0f68968f19c5821f629a9a1ed33

Observation 7e502ea8-cd4f-44bb-a91c-8aacd2cf5bc5 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Scaling Inference-Efficient Language Models Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.279768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.279768Z digest=sha256:f504815f8fad5a961f129ee43a94c4d7b77014cfec822507a6687b18fbbf82b6

Observation 00bc4bce-51e1-4033-a47d-bff38ca9befd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Scaling Inference-Efficient Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.317998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.317998Z digest=sha256:1c4d679c0d250f781bcc40960da5bf3689fe4c028b9f376a2d6d1bd743fd3d5b

Observation f36b0fb5-36bf-46ce-bafd-12008f918a7d · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Scaling Inference-Efficient Language Models OPT: Open Pre-trained Transformer Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.358954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.358954Z digest=sha256:078d79d49ed8c1d5182de04e364e6d606c740c67f46820510a5b2d34b1c3782f

Observation 9809848f-786e-40cc-b902-b7d36b2a17df · outbound

This paper cites (Center) We indicate the relationship between inference latency and hidden size with the number of layers fixed.

Scaling Inference-Efficient Language Models (Center) We indicate the relationship between inference latency and hidden size with the number of layers fixed

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.471291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.557585Z digest=sha256:0a175aeec0546b16610410e37fdcc14b0c87730ad5cc15950e46d7143d919340

Observation 92eb9d33-d361-4bba-9201-2c641ead1be5 · outbound

This paper cites The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024).

Scaling Inference-Efficient Language Models The evaluated models include LLaMA (Touvron et al., 2023a), Qwen (Yang et al., 2024), Gemma (Team et al., 2024a;b), and MiniCPM (Hu et al., 2024)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:43:30.400895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.564505Z digest=sha256:06b366d263ea3812ddc389812e58377800cbe58b45baab239655a0436e026c67

Observation c0cfc69b-3642-40db-ac9d-47249834c10f · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.324478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.567890Z digest=sha256:4f74bea5f08698dc5135a832ed9af148a2ee2c9684b17394831a6abfa6d0de8d

Observation 8b000cb5-25d9-4ce2-bc4c-40069c0d50c6 · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.178497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.571468Z digest=sha256:4e431f34ec4e1f1dad4f56d5aa72518a541393eca42efd3cb1564d893f0968ac

Observation a6da3976-01ab-456d-9366-3039c7b87aaa · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.134727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.574447Z digest=sha256:de9b19af2a7a7d66e90c5b6f9fe16c34e983854ae7077adccaf6bed7ee182c7e

Observation 459ddbfe-1a6c-4f01-bc00-a4463a9c4683 · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.123552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.577850Z digest=sha256:b8f04c3dfeefa5197d90a56ccdbabb2ffc4b96110200ff30afd7683735d35251

Observation e805e0bf-72d8-4660-b255-d504e1315cfe · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.496911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.537141Z digest=sha256:6fe91e46a0e13f09c9dc5abd38ce60b6bb805e5ce27ee4d66cad86998d5c4c0d

Observation 6266e32d-d3a6-4a57-b7b7-76549fca7d7b · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Scaling Inference-Efficient Language Models Efficient Streaming Language Models with Attention Sinks

Reference 1921

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.256420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.256420Z digest=sha256:44255d31214efd066a15ea1cbd68e916caf179dd5bb1b675225b428adc90fe02

Observation ef6e98a8-048f-41af-abf5-fa1949c0175f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Scaling Inference-Efficient Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 1961

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.198924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.198924Z digest=sha256:523b922808b289e521d1e03dbeb77c8e33702d2b7076dce5e0a023a7d1be062a

Observation 72478f8f-797d-4a77-8054-4166b6a683fe · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Scaling Inference-Efficient Language Models Observational Scaling Laws and the Predictability of Language Model Performance

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.116548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.116548Z digest=sha256:f434445c6a73901740cd949d4cc042987cbd730b60b9505727535eedb3b78ec9

Observation 5f881308-29a5-4563-ba0e-b23af1b6e46f · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Scaling Inference-Efficient Language Models Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.080074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.080074Z digest=sha256:f2196d3ab874dbf95a014fff5573c9ef8119959c0b096a63a3b47bda3857346b

Observation f4d96b4e-56c0-4bf1-98d8-0e3684f6fc53 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Scaling Inference-Efficient Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.016527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.016527Z digest=sha256:7b83b7c8ec3c6c41cb557afe41bd976db10833072915fa06684d1b14a45d08af

Observation c9fc260f-b242-46db-99aa-ed15d558a6b6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Inference-Efficient Language Models Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.768781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.768781Z digest=sha256:5b066c812c9b34414a728b44e07f62a85141d2eb11ae4042b01b535c853869d1

Observation e3130bf6-75bc-4e61-824d-6417abde6ec9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Inference-Efficient Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.672091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.672091Z digest=sha256:1eb5d96132a95d5f6321a7ec5404e63e742f174d4e607113301340bff56514a9

Observation 9fce7d57-3670-412a-b473-88a41959fe5e · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Scaling Inference-Efficient Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.612128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.612128Z digest=sha256:10529f79f8f8b86eed1d79095dfe1af454fe02bda89344e039d0166c77a51cb9

Observation f9b70a2b-b9ac-4d21-ad99-c0df5c2c6344 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Inference-Efficient Language Models The Llama 3 Herd of Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.819084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.819084Z digest=sha256:b61e7c1977abd0f63854bd7f1a9ece69a24f28362cda7262bc6b0705f9a99813

Observation 764af2e5-a84a-46ec-ab84-f40ab30dc6e6 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Scaling Inference-Efficient Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.616404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.616404Z digest=sha256:c802adf942fa43b61e40ed98f24d9b8ac3e13f85c32a06f8bbbbe80b780a0dd0

Observation 2a3c5454-eb55-4b54-8778-c051d3ff73bf · outbound

This paper cites Qwen Technical Report.

Scaling Inference-Efficient Language Models Qwen Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.607548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.607548Z digest=sha256:92b7e99cb751ea5bd9ac96a6d6d96e23bf8658707c2775e57da922cf5a1d9024

Observation 615efc87-a32f-4aea-8c20-8bfd1f41ff8c · outbound

This paper cites Phi-4 Technical Report.

Scaling Inference-Efficient Language Models Phi-4 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.585716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.585716Z digest=sha256:96acfaa3afab464b5a5acaf3e3697e8c2ef02022bb4efa4399af5a25cc98fe64

Observation 9965cdfd-08b6-47ba-9256-79a8b6f2a577 · outbound

This paper cites CHAI: Clustered Head Attention for Efficient LLM Inference.

Scaling Inference-Efficient Language Models CHAI: Clustered Head Attention for Efficient LLM Inference

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:43:30.086880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:28.601962Z digest=sha256:0bbb0319d4921e6698d066a6baa595a7bf006ee5977a44c1c67d18aff5bc5455

Observation 8e4fcd9c-f716-4138-b306-c9f8e6b8b30e · outbound

This paper cites an unresolved cited work.

Scaling Inference-Efficient Language Models Unresolved cited work

Reference 2048

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:43:30.533224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T00:43:29.385782Z digest=sha256:75693c3f9b2a15ecdb42ca70eb14e5c28f1833e887f1ed96c13863573cd5c9ed

Pith citing papers

Observation 49760f12-0291-4357-a221-0caf1b0379f9 · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Scaling Inference-Efficient Language Models

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:40.215851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:40.215851Z digest=sha256:66d0c43977eccd71804a6b4fd89b5b6d753f98e97b3364d93c4d2f5dc632f968

Observation b109582b-e94e-4858-853b-129ad6b632fc · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Inference-Efficient Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.965606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:303227b1c6ddfab50df9035a9e8d5d66036bc2fd3019e21fb1ce3ded40637495

Observation ad67b25a-8e8c-41f2-9975-9de5ae1c9a41 · inbound

Comprehensive AI governance requires addressing non-model gains cites this paper.

Comprehensive AI governance requires addressing non-model gains Scaling Inference-Efficient Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:31.619579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T08:11:23.860723Z digest=sha256:e7cf530bbe322058a5533e9591ac61a0b6aed07d4820c028b545e417c4a34534