Pith. sign in

Paper Citation Record · LEDGER

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2312.11514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11514 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:40:01.304833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:29:50.698429Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fc09583-7952-4354-8218-5fc146502e6d · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.814649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:686cf731fe72a94802788b55c7a235c50f8ebf19caf2b069ece0500d350ff634

Observation 4a1720d8-572a-4437-9210-a5837ea39253 · inbound

Less is More: Optimizing Function Calling for LLM Execution on Edge Devices cites this paper.

Less is More: Optimizing Function Calling for LLM Execution on Edge Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:11.743252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:11.743252Z digest=sha256:f485a40f651027d9e5aeb1559f4a02ad73e3850b1990ccf1a386dec0d878a3ba

Observation e51fb764-e5d8-40c2-a466-1724c0b938b0 · inbound

Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems cites this paper.

Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:05:57.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:05:57.594324Z digest=sha256:c1b6146f567bbf349136df6b63ba0c74957f26d6934025a56e59b001bd14b5b0

Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · inbound

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking cites this paper.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.880444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.880444Z digest=sha256:d21fa2174dbde38e95124d874ed8f0be104c1eb4dc6beb707ccbe4d58b2ecdf7

Observation 0f94a630-6bbb-421f-aaf1-0c5ef7562a01 · inbound

Mixture of Hidden-Dimensions Transformer cites this paper.

Mixture of Hidden-Dimensions Transformer LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.416957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.416957Z digest=sha256:1ad20c64c5d5aa33c82650c0d6a49c52e850aa2c85606f0f6467bb18dff06c9b

Observation b14fe271-994e-4c4a-ae18-d8aff161c048 · inbound

Post-Training Statistical Calibration for Higher Activation Sparsity cites this paper.

Post-Training Statistical Calibration for Higher Activation Sparsity LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:10:12.264153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:10:12.264153Z digest=sha256:db40e49ee1cbb2a9a0261dffda6b8e916cc457580fb5e70bf56dad1935b1cf14

Observation 834f1932-cb9d-4cf3-a64b-1056f4291f18 · inbound

Creating an LLM-based AI-agent: A high-level methodology towards enhancing LLMs with APIs cites this paper.

Creating an LLM-based AI-agent: A high-level methodology towards enhancing LLMs with APIs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:04.423298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:38:04.423298Z digest=sha256:eb3433fdda740af957d5a7186d7efbe93b4df18d812e263848b14619c0d16da1

Observation 2a5fadc7-b381-441b-b00e-acea44634311 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.128373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.128373Z digest=sha256:b7c028515b0479b5858f387687d16c9c1172147ce0bdcde7204857c2a12c2703

Observation 6e8de6fe-2d1e-4370-b559-347890025dc0 · inbound

Accelerating Retrieval-Augmented Generation cites this paper.

Accelerating Retrieval-Augmented Generation LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:22.972992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:22.972992Z digest=sha256:de51d88a9a350e57eda63437e78c7b137ebcf70dac6aae843e997fca6d0178e3

Observation 59c233a4-120d-470f-b98c-b007632f82a6 · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.059448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.059448Z digest=sha256:2baa71aa0c75c7d8445ab50474793ddbf5cc3b3e81f173f5575caad36721df3b

Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.738166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.738166Z digest=sha256:aff85a303972c48e25893b59ddc30c45bb064aa26341700d09a880a1be3a04bc

Observation ff554477-e8fb-4aa2-b6a8-5dbdbeff8b8f · inbound

Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning cites this paper.

Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:11:01.077187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:11:01.077187Z digest=sha256:b1782bee4e07ee838e3f4fbe7e3b4d5e4d95d375e959ddd8ea13ab88699b437c

Observation 363fa428-9e58-4982-93c0-61d1468f4cde · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.201744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.201744Z digest=sha256:b8dccc807ea1fd2f289ed54acdcacc115da247bd39f82c26e4ed882801141419

Observation a21b22ec-6027-4945-98ae-68596982659f · inbound

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge cites this paper.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.786243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.786243Z digest=sha256:391f4bdd69b64727c8ffd5df807175dfb23a51d7068c709d1464168c2138c8b5

Observation 999649e1-e61e-4631-9e97-150b5e4bd1ea · inbound

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices cites this paper.

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:37.615578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:37.615578Z digest=sha256:f83c78301662b97609d52698d0abf5f621574c731fae5e431ffbf9673d9a846b

Observation 0d2a9076-757c-42be-9773-19f5cd28eba9 · inbound

MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection cites this paper.

MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:40:01.304833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:40:01.304833Z digest=sha256:16afb6bf2fa2bc7fb76abebc8fd592e31c00e772eaf29790c5ffad3dda64a9f0

Observation ec61dee9-143f-4453-b8f1-6702b357150e · inbound

Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs cites this paper.

Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:41:55.943745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:41:55.943745Z digest=sha256:91e5bb7a8d74765bec0391959029ebc16f52bc519a2ad9f5cd16d687073896ae

Observation 7c460a86-acb3-40fa-85a6-8eafae05757b · inbound

AVEC: Bootstrapping Privacy for Local LLMs cites this paper.

AVEC: Bootstrapping Privacy for Local LLMs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:47:45.635928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:47:45.635928Z digest=sha256:60f4e51d3dc2ab5272998a02a9306dd25b25b5589272c48c6053db311d7e3544

Observation 896e74e0-c3b5-46cd-93fe-e3973c893988 · inbound

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms cites this paper.

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.239648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:36:26.337511Z digest=sha256:4ae3fb32ef0cdf6b3db2c161863f2b68e0312ca1d348a995924b56c424774303

Observation 85f463d2-bd83-4176-971f-db403139b2c7 · inbound

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories cites this paper.

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:13.449101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T08:11:57.929452Z digest=sha256:f8d0b7629a5baa43b221c23ed627bc4c545f9498ffbe4b4dd02656a13931fdfa

Observation f216b5e0-9a4a-4794-a5c0-b92ed0323eb7 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.573557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:0d6246788893280f57d0a4cabf455b8fc34d6232bc48e24b52518e7c6547979a

Observation 99a6258b-ad0c-4ee1-b3eb-a4de5b1f90f0 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:21bfb3d2f45f755319c66049912f2b8ef901048dfbb8d273b18d0e2ccc30069f

Observation bc1102b2-94f2-4125-914d-f382c280d540 · inbound

EnerInfer: Energy-Aware On-Device LLM Inference cites this paper.

EnerInfer: Energy-Aware On-Device LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:29:50.700025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T07:58:29.150228Z digest=sha256:71286fe375c065a48de12400167f5137da0d9df93bac131ad6ac92cc2b2725ba

Observation b042df67-8388-4494-b6f0-8903e4813f16 · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:58.904402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:58.904402Z digest=sha256:a406e7ce3bf502c2185d17af1f29cf7ddd5a5807a1abb2716eac256b932b053c

Observation 7b3eedeb-cef1-4533-871f-cac10850f96a · inbound

A CXL Memory Rack for Multi-Turn LLM Serving cites this paper.

A CXL Memory Rack for Multi-Turn LLM Serving LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:58:24.558673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:58:24.558673Z digest=sha256:6c4ec64eb54e70b44a049427670e01d7360bc294c7451859453268527a57a30b

Observation 3cb3c67c-0732-4b33-b3fe-af6857632992 · inbound

Architectural Implications of Agentic AI Workflows cites this paper.

Architectural Implications of Agentic AI Workflows LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.501057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:12.501057Z digest=sha256:35a202c3a3874b1512cdc929453ba7dd5ac66b9071f7f1f679a6d69336980582

Observation 79336d5a-5211-4888-896e-006ed9be5693 · inbound

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI cites this paper.

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T14:26:28.950463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:26:28.950463Z digest=sha256:c143862c38f1fb54312a87c6745ae15569b68124eafe81c211accc851e424b3f

Observation 868a362e-067a-4b9b-8e00-608c3202bbf5 · inbound

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI cites this paper.

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:52.862515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:17:52.862515Z digest=sha256:a20e7819884374cbaea1099cb63f12f211d25caaa949d28d747db5a193f34174