Pith. sign in

Paper Citation Record · LEDGER

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2312.11514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11514 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:26:28.950463Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:29:50.698429Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fc09583-7952-4354-8218-5fc146502e6d · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.814649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:2cb8d667fa44440eb3fccd9c59093d63555c3062d0b1c7621169bb918326b0a0

Observation 4a1720d8-572a-4437-9210-a5837ea39253 · inbound

Less is More: Optimizing Function Calling for LLM Execution on Edge Devices cites this paper.

Less is More: Optimizing Function Calling for LLM Execution on Edge Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:11.743252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:11.743252Z digest=sha256:f485a40f651027d9e5aeb1559f4a02ad73e3850b1990ccf1a386dec0d878a3ba

Observation e51fb764-e5d8-40c2-a466-1724c0b938b0 · inbound

Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems cites this paper.

Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:05:57.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:05:57.594324Z digest=sha256:c1b6146f567bbf349136df6b63ba0c74957f26d6934025a56e59b001bd14b5b0

Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · inbound

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking cites this paper.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.880444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.880444Z digest=sha256:d94ca1ffaccc809618b62129ef33103bfdd94cc00cb27c7b1906342a0faa536c

Observation 0f94a630-6bbb-421f-aaf1-0c5ef7562a01 · inbound

Mixture of Hidden-Dimensions Transformer cites this paper.

Mixture of Hidden-Dimensions Transformer LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.416957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.416957Z digest=sha256:a6d9ffd9abc21c9c3b35e868a499d0828638cc2aa01a776a6c35664934c5d870

Observation b14fe271-994e-4c4a-ae18-d8aff161c048 · inbound

Post-Training Statistical Calibration for Higher Activation Sparsity cites this paper.

Post-Training Statistical Calibration for Higher Activation Sparsity LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:10:12.264153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:10:12.264153Z digest=sha256:db40e49ee1cbb2a9a0261dffda6b8e916cc457580fb5e70bf56dad1935b1cf14

Observation 834f1932-cb9d-4cf3-a64b-1056f4291f18 · inbound

Creating an LLM-based AI-agent: A high-level methodology towards enhancing LLMs with APIs cites this paper.

Creating an LLM-based AI-agent: A high-level methodology towards enhancing LLMs with APIs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:04.423298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:38:04.423298Z digest=sha256:eb3433fdda740af957d5a7186d7efbe93b4df18d812e263848b14619c0d16da1

Observation 2a5fadc7-b381-441b-b00e-acea44634311 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.128373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.128373Z digest=sha256:b7c028515b0479b5858f387687d16c9c1172147ce0bdcde7204857c2a12c2703

Observation 6e8de6fe-2d1e-4370-b559-347890025dc0 · inbound

Accelerating Retrieval-Augmented Generation cites this paper.

Accelerating Retrieval-Augmented Generation LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:22.972992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:22.972992Z digest=sha256:449dfbdd104db963623659fbb94d9b9e4261c3556bbffe343aa9755e19fb7988

Observation 59c233a4-120d-470f-b98c-b007632f82a6 · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.059448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.059448Z digest=sha256:88f767cd9ba82d4868f9d3a4668cf4ed8a7ce5fc9508815b29df4ec8aa8ae5b8

Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.738166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.738166Z digest=sha256:aff85a303972c48e25893b59ddc30c45bb064aa26341700d09a880a1be3a04bc

Observation ff554477-e8fb-4aa2-b6a8-5dbdbeff8b8f · inbound

Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning cites this paper.

Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:11:01.077187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:11:01.077187Z digest=sha256:b1782bee4e07ee838e3f4fbe7e3b4d5e4d95d375e959ddd8ea13ab88699b437c

Observation 363fa428-9e58-4982-93c0-61d1468f4cde · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.201744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.201744Z digest=sha256:140219198881d8fe50a09c87b955dd76bb330578af81475b27af031a4ba6714e

Observation a21b22ec-6027-4945-98ae-68596982659f · inbound

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge cites this paper.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.786243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.786243Z digest=sha256:391f4bdd69b64727c8ffd5df807175dfb23a51d7068c709d1464168c2138c8b5

Observation 999649e1-e61e-4631-9e97-150b5e4bd1ea · inbound

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices cites this paper.

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:37.615578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:37.615578Z digest=sha256:1a9ee2600d563994681e809ad353002e4ba6cbaf00d9b34d1b80f02e7d90059f

Observation ec61dee9-143f-4453-b8f1-6702b357150e · inbound

Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs cites this paper.

Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:41:55.943745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:41:55.943745Z digest=sha256:25cacd592bed4a4899559e794139e9d7022d5da647680435954656091cb9b892

Observation 7c460a86-acb3-40fa-85a6-8eafae05757b · inbound

AVEC: Bootstrapping Privacy for Local LLMs cites this paper.

AVEC: Bootstrapping Privacy for Local LLMs LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:47:45.635928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:47:45.635928Z digest=sha256:60f4e51d3dc2ab5272998a02a9306dd25b25b5589272c48c6053db311d7e3544

Observation 896e74e0-c3b5-46cd-93fe-e3973c893988 · inbound

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms cites this paper.

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.239648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:36:26.337511Z digest=sha256:3959d39efe037fa8d55f8c8a25e77fd064cffd71ab7db2c61891523440962a3a

Observation 85f463d2-bd83-4176-971f-db403139b2c7 · inbound

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories cites this paper.

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:13.449101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T08:11:57.929452Z digest=sha256:1d96008813c2cb0865a7894d808ee74138fec4601f01c4522d2c4e335cb2b325

Observation f216b5e0-9a4a-4794-a5c0-b92ed0323eb7 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.573557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:2bae292b68494669b043183d11be2317c624bc86bd8a84832df53e731e6b4709

Observation 99a6258b-ad0c-4ee1-b3eb-a4de5b1f90f0 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:4d668ac8e5114e338a37409b66d40db44bb03f0c4589ca10138da115bc03f5a1

Observation bc1102b2-94f2-4125-914d-f382c280d540 · inbound

EnerInfer: Energy-Aware On-Device LLM Inference cites this paper.

EnerInfer: Energy-Aware On-Device LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:29:50.700025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T07:58:29.150228Z digest=sha256:f348c6e5efe92bcaf1f47f502deaa067b78535b40c072b4a0a8e93412b51460b

Observation b042df67-8388-4494-b6f0-8903e4813f16 · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:58.904402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:58.904402Z digest=sha256:b4be4f8de10f4e7695c7886e0077627d43b0d7e8c1c9d143e450c61fe217ef4e

Observation 7b3eedeb-cef1-4533-871f-cac10850f96a · inbound

A CXL Memory Rack for Multi-Turn LLM Serving cites this paper.

A CXL Memory Rack for Multi-Turn LLM Serving LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:58:24.558673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:58:24.558673Z digest=sha256:b4b71d0bf93c52cfe4b8763f7c786fa90a34365c282633cc60020dbf07054e0c

Observation 3cb3c67c-0732-4b33-b3fe-af6857632992 · inbound

Architectural Implications of Agentic AI Workflows cites this paper.

Architectural Implications of Agentic AI Workflows LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.501057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:12.501057Z digest=sha256:35a202c3a3874b1512cdc929453ba7dd5ac66b9071f7f1f679a6d69336980582

Observation 79336d5a-5211-4888-896e-006ed9be5693 · inbound

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI cites this paper.

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T14:26:28.950463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:26:28.950463Z digest=sha256:07128bb2a47eb05342eb659dc702ec62f3275d97d0b9e7a6e74db7aed2a1762c