Pith. sign in

Paper Citation Record · LEDGER

Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2309.10285.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10285 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.672927Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 72d3a5b1-ca63-4a16-bac2-3fb80da30ac2 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.672927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.672927Z digest=sha256:f5643aa7df529e75de950dc068611c6949b7fc47755ca9ed3c885e761fe61caa

Observation 0664cf4f-2f0f-4f70-805c-12afad15f361 · inbound

RAP: Runtime Adaptive Pruning for LLM Inference cites this paper.

RAP: Runtime Adaptive Pruning for LLM Inference Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:21:35.642487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:20:41.739571Z digest=sha256:31a76b3e21c0b1f0760199c5af43497138e1fcec36e33ff0c0881cf0f1282c39

Observation a4826869-d6d3-49e4-adbe-945c7e9bd8b8 · inbound

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs cites this paper.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.552122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:54367a6f7e8ea200cb894eee6e6ca610900255ac1ecf63565181dc4ffe81b48b

Observation f690e976-331a-45ab-a40b-0650d905e719 · inbound

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models cites this paper.

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T11:47:19.007455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:47:19.007455Z digest=sha256:1155189b13171391094f7f5a009535d2fa28f2f37f8ccc332e51a42b21a2213b

Observation caea12ae-603b-4c34-a5ee-11c1223ddaf5 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.309428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:5d40d503b19555ad560623b3c070dc68ac15dc055976ac2ad76f1d5453ad6d06

Observation c139bd58-344c-4240-af5e-218b7f1200cc · inbound

ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs cites this paper.

ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T19:37:43.766024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T19:35:51.992891Z digest=sha256:ef6a78d826e25cc84a9dbcb424a65e0627147af09933c662c7ffa6236531cf06

Observation cdefe15b-5862-4af1-942a-d27eade948f5 · inbound

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy cites this paper.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.718176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:7cbd1980cf5a6f0cb03de27bc4b9deab4da0a410964feb1b8b7a01ab21dc1543

Observation 3b184a85-9f3c-4509-a578-95e4ef8985b9 · inbound

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference cites this paper.

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:09:28.755271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:09:28.755271Z digest=sha256:2d3551995e49914b4403c06cdc8c61789d15f8cb4617a21c33ec2870a54cd178