Pith. sign in

Paper Citation Record · LEDGER

Inference economics of language models

As of 19 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 13 inbound Pith citation observations for arXiv:2506.04645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04645 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:48.474937Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:21:29.458359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.165925Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a6db3ea-157f-4845-9f1e-2529f6abccbe · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Inference economics of language models PaLM: Scaling Language Modeling with Pathways

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.385356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.385356Z digest=sha256:4ad529c6b0e49f162cb7a48c5785ff1bc0a192ee1efcac8275a20485cba2bdf1

Observation 35eb7c2d-c09b-4705-ad24-12896af4fd21 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Inference economics of language models Fast Inference from Transformers via Speculative Decoding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.474937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.474937Z digest=sha256:a12ed507329d6c1ee8f2693fb7ca325a10d496e9507b423535b4d4343cbf1b94

Pith citing papers

Observation 2e1dac95-2868-46bb-aa47-e2a6dd506406 · inbound

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches cites this paper.

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:41:25.963890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T13:36:55.938673Z digest=sha256:2bdb693a26a8b6e09406802eeaa9c3a212087884ec15d42242847f1ea31cb822

Observation a72bb78f-99c1-4ada-aa8e-fcb66231a207 · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:43:32.112169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:41:18.795730Z digest=sha256:831e7a616a127f4ec7ed149e6cddcee49d237a28dc6659e51fc55fe7de99be89

Observation f2329126-04e4-4dd8-9b45-7dde365fa7df · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.200797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T16:54:58.406395Z digest=sha256:625ba7f8593603b94fd7d67b0c8e578114b2f067b4819eb27e16458794e8589e

Observation d9dea4af-5767-436c-aa02-d6ee6dc4ae9a · inbound

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods cites this paper.

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods Inference economics of language models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.423669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:37:49.558450Z digest=sha256:6695f9e0b50063c5c5c056ee89c63f44b13be447fe23b752efc1bfe7a5e76b6b

Observation e8c8a6e4-0338-4c8e-8f8e-2b827045b8b8 · inbound

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants cites this paper.

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Inference economics of language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.158798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:34:16.075235Z digest=sha256:212db5f16f3a379fa960c6966d72b934579cca350d0a28f4bc8ed8b0da038bb7

Observation 97493b5c-aa08-4af4-81b6-0f5dcdc1aa20 · inbound

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development cites this paper.

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development Inference economics of language models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:12.070833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T18:29:18.485336Z digest=sha256:2c44a01cedaeaf1f19c08d9f004fb4fa65044d260f6615042533d4ed6d55ade3

Observation dd6cd589-49bd-437a-93f3-2fc4e4b46384 · inbound

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference cites this paper.

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference Inference economics of language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:24.637089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T12:36:23.320194Z digest=sha256:4f84f61695168a7aea66099d870d452a19732141ed157ef6d34a926405b40c03

Observation 1adaf6bc-9594-4451-8151-2bf8b9e2e8e5 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.782979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:0fe25f68ef610bbd34c08da16823ca9dfef7198c44c81d4851c3f19480746301

Observation 122dca54-a96d-4745-ba2a-23dba8e89d7b · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.141713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:55d4490257c1de2a7b9aa67d8292458ac3314919c5fe1504a7adec8cfadf5b17

Observation d8fd2f96-9403-41d8-a630-009a86a578d4 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.196297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:51c8898b6edf7564c7e2e2d3324ac30616362da024060ea1cde9adb71e060fb8

Observation 987461ec-5ba4-481e-bfc2-a79665687c7f · inbound

Efficient Clustering with Provable Guardrails for LLM Inference at Scale cites this paper.

Efficient Clustering with Provable Guardrails for LLM Inference at Scale Inference economics of language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:03:18.785055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:03:18.785055Z digest=sha256:cb39589291a80bf9dc7f1f06d05d6b8b107e90840038fa314e47011e227dcf8f

Observation c8dc26db-bf8f-4bb6-b57e-7126535a1c92 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Inference economics of language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.155145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.155145Z digest=sha256:ea33320fe0666b13c56558197e97d1184ea062d0149234e2df841e88089c32df

Observation 9931379b-adf8-4bd2-87e6-95d5d15fe743 · inbound

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems cites this paper.

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems Inference economics of language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:21:29.458359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:21:29.458359Z digest=sha256:3b7898f5079f4dec7319522fed46f5da160a8fd0067753977441610174f85347