Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:30:36.140273Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2412.01380.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:30:36.140273Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:15:49.289390Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T13:21:35.626449Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99550ca1-e45d-4b67-b230-96d096becd2c · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9407622e-fc95-4455-b73a-5bfdd1865fa4 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f6ffc3-c380-4cef-9e91-386af7bbcd0f · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1078dc-490d-4158-8169-e5a88d67312c · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 845a17a7-00dc-4003-81bc-f2a6868c3fdb · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Smartphones beat dram drum to meet performance demand, 2021
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c977e7a0-181c-496e-9c1d-4172fe11b670 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking How much ram should a phone have, 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6f33b5aa-e1b5-4aa1-91f2-422074d7cdd6 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking N., Fan, A., Auli, M., and Grangier, D
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d12ab3-7d07-47d2-832b-b0e20c642ee6 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa89924d-8db7-4aa8-8592-bfca4782e7e7 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c840b7-2ac2-4992-bc00-60d1e7fde47f · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Code artifact for efficient llm inference using dynamic input pruning and cache-aware masking, March 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c96977b5-55f0-4e86-adc9-c0a463f85676 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Alistarh, D
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a7f179-82da-4fc0-bfa6-429a915aa293 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c69a890-7a6f-436b-a286-0c27c9a20853 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A framework for few-shot language model evaluation, 07 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a84b3fb-b85c-4189-829e-062bea9a86f5 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking W., and Keutzer, K
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a52cae37-1f11-4309-9d1f-2f0678d8cf28 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lorenz, J
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation de89ce20-02e1-44f0-b1f1-3b77bbc824a1 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Measuring Massive Multitask Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c159906b-805b-451f-8773-775215900693 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980c6f8d-38d7-4218-89d8-32e703c6c764 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bcb492-48e7-4f5d-9a59-c9ddd76ae092 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lin, C
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7b972b95-ad44-4721-912b-66af0df3758f · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Challenges and trends of sram-based computing-in-memory for ai edge devices
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1ffd6fdc-b3bc-4870-bc4f-035d3326d324 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mistral 7B
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b47f41-f004-43be-98c0-9698e15ca0fc · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: Which is better? Advances in neural information processing systems, 36: 0 62414--62427, 2023
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f0f7afbc-d75c-482f-8ae6-ecfe48b82a42 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: which is better? Advances in neural information processing systems, 36, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 347859e4-90cd-417d-9aaf-18d7d12244c1 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking H., Gonzalez, J., Zhang, H., and Stoica, I
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1c6410-ba49-4e31-994f-f5ca55b5f7af · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple nvme vs ufs 4.0 storage, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5d5c0d83-edb6-4a88-b3f4-a6c401ad461e · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a047825c-5268-4438-aacf-4ff4eee851a3 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea440abe-924e-4ad6-9b97-e15a61138d6f · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Deja vu: Contextual sparsity for efficient llms at inference time
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c17f66-d939-4db3-a385-2538d8109a1c · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Vassilvitskii, S
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661d8d82-7e10-4393-8148-858695803850 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llm-pruner: On the structural pruning of large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33594af2-a8c9-4ee6-8f9b-88a2eb90954b · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c22a46-6e6f-4636-979a-6258aa59bcf4 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A White Paper on Neural Network Quantization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77168a8e-e847-4c47-bc55-aa7753c693a5 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Comprehensive Overview of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87065d3-03dc-4e94-aa3c-bf575b49cace · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llama 2: Early adopters’ utilization of meta’s new open-source pretrained model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0808ae4c-4854-4cf7-815b-413bff1b881e · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b564d181-3870-4d71-91a6-bba295237c1c · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Effective mimicry of belady’s min policy
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1f17811-df1d-40ef-b74a-c28de383297f · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking R., Hestness, J., and Dey, N
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9653c257-04d8-4d09-be4a-8622a7579b98 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0727e386-f3a9-497b-8cf2-561e39e4bada · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba622ce4-2691-4519-87dd-712870577f6a · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb59a358-0219-40a4-8d81-8de302534ebf · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sparse language models with relu activations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a0031b78-b421-4dd9-8280-2c322a9c32ae · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Simple and Effective Pruning Approach for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3088ac87-2ccb-4896-9e5c-2459bfbdc331 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTVQ: The Blessing of Dimensionality for LLM Quantization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6655a9d5-9afc-413a-8fd7-e7ce627e7381 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The LLM Surgeon
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170d6ec3-2d9b-4302-848e-39d0cc1729f8 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9bc7e5cb-0308-46eb-8502-a7c4732ba1c3 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple silicon --- Wikipedia , the free encyclopedia, 2024 a
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 33a9ae10-e9ca-4ed8-98e0-17fd196aa542 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking List of qualcomm snapdragon systems on chips --- Wikipedia , the free encyclopedia, 2024 b
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3318d1f0-5035-4b36-9cf9-243395fd3c4e · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Flash memory --- Wikipedia , the free encyclopedia, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8c7d2330-15c8-4845-b7b3-f9906a84d1b9 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c025d0-cc7a-4893-b154-3e362a848f6d · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f10859-551d-4b20-81e3-9e7caedddce9 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OPT: Open Pre-trained Transformer Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667f0a38-94ac-477e-93b1-340159c08a00 · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Survey of Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d4b92d-51ca-4b6d-b01c-ea9de18e957e · outbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e8b28b-7346-4cbf-9aca-f48e6a169591 · inbound
RAP: Runtime Adaptive Pruning for LLM Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 244a3fa0-3f99-483d-b108-e6139357ad16 · inbound
Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1490825-f9dd-4591-a84c-a20a207d8aab · inbound
AutoNeural: Co-Designing Vision-Language Models for NPU Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.