Pith. sign in

Paper Citation Record · LEDGER

Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.01698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01698 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:22.757378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:50.618032Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bd779b54-fef4-4499-b02c-d0d4df4e3041 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.757378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.757378Z digest=sha256:af899523f40379340c0280fecc5227e30f78f49e1b01b4b8197df65314a04714

Observation e72a6acf-d675-41d8-8f5c-aa33806f719b · inbound

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference cites this paper.

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:17:08.008423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:16:31.655330Z digest=sha256:99793417a362c267ed0e63d82b83a0c358f00ecb3e69ed4edaed4e1c03e6a494

Observation fd25a10f-cd5d-43a4-96c0-d32870f5f889 · inbound

Scaling Intelligence: Designing Data Centers for Next-Gen Language Models cites this paper.

Scaling Intelligence: Designing Data Centers for Next-Gen Language Models Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:08.834409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:08.834409Z digest=sha256:4369200e8c4f8caea5f4c6db22bd1b1d7da634c184a76228d8a945238dc5648c

Observation 5da1ad3d-e9c4-4aa5-927c-c81c802741e6 · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:27.508377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T18:43:06.803528Z digest=sha256:d7775e75f3ba7005572bed99b48df6de4581b782342b544d4b65946e069bb2ec

Observation a5d85336-3a48-40c7-9833-bbd5d7ed03a1 · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T15:09:42.841402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:09:42.841402Z digest=sha256:6fb19a73cca018bf87b6d350cb70ed56d53886600e5f4998dae01044a0790972

Observation 2d73e722-a1b4-46b1-b429-697183779ab8 · inbound

SOLAR: AI-Powered Speed-of-Light Performance Analysis cites this paper.

SOLAR: AI-Powered Speed-of-Light Performance Analysis Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.769879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T01:30:29.503481Z digest=sha256:c5e9f98d474b08c4a83f660cc79d938de29b2fda4b2e847ecda9acdc787dc4cc

Observation f1cfe1ea-0cbd-4e32-ab63-b576b25c3c1e · inbound

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems cites this paper.

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-12T11:05:56.233115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T11:05:56.233115Z digest=sha256:0879583a4e21bd88101ce4669e392277730242cf4eb52f8b556fa2a331e52d12

Observation 6f22cb1a-6aeb-4615-b333-8e06171bedab · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.131906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:39ad818fb88b537098bfda2e6056bfae67ee5087863f126a091f69fe5ffd5c23

Observation 4e5a5394-2bf8-4750-b027-16d66ad1e45b · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:50.645437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:ef617c9edc88487c89014f2bab3196d7214af504049ba0ee4a06d90f0a309803

Observation c51afc1d-95f4-4a48-b5dd-c68499f93faf · inbound

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters cites this paper.

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T04:49:44.520794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:49:44.520794Z digest=sha256:fce5c9479b4a7f80afb7c47f717d9779412b3e927b6fa5ac3b33f9447d8a87ac

Observation 6ecb6663-a376-4b03-b5ae-600d7647461c · inbound

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving cites this paper.

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:42:53.666900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:42:53.666900Z digest=sha256:c9908f15b6106ab1cf7a9b0c1977407f010f0cda20d440bc1a5ea2dba0687787

Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · inbound

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving cites this paper.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.136193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.136193Z digest=sha256:7bc77a8168ddf0cef7008a317fa0332726f1e8b71bcd56cce48f636c6b8db044

Observation c4b1a9e1-72b9-4d48-981d-c4f5cd7b2c87 · inbound

RAG-Stack: Co-Optimizing RAG Serving Performance and Quality cites this paper.

RAG-Stack: Co-Optimizing RAG Serving Performance and Quality Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T18:13:12.603227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:13:12.603227Z digest=sha256:ffe4f90043d0b963c1f52d98fa81ec062c5920ced0b99848a682d24a429ddca9