Pith. sign in

Paper Citation Record · LEDGER

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2412.01380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01380 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:30:36.140273Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:15:49.289390Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:21:35.626449Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99550ca1-e45d-4b67-b230-96d096becd2c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.868916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.868916Z digest=sha256:0f2fac58f6b65501df46f8d186b3978529c484d3b171dbd25ce1c108d9b2eae3

Observation 9407622e-fc95-4455-b73a-5bfdd1865fa4 · outbound

This paper cites ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.874693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.874693Z digest=sha256:53d336ee4715748537e60531e62ed413592d38c1fcbd63cef97c88e4ff821f0b

Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.880444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.880444Z digest=sha256:a883ea6f6bd49c39c72ba403aaa3905bc6aa6de7cc266d4acc918fe348d83d67

Observation 54f6ffc3-c380-4cef-9e91-386af7bbcd0f · outbound

This paper cites Qwen Technical Report.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.885851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.885851Z digest=sha256:2b5be3059eb583944d284d137c8e0500ac17b7fe52e002dca59170d42faa9292

Observation cc1078dc-490d-4158-8169-e5a88d67312c · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:37.255188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.890455Z digest=sha256:b83d66678c5f627c3ad8ea33e2a2372a9b3ba0ef10864c9cd2317a22116247c6

Observation 845a17a7-00dc-4003-81bc-f2a6868c3fdb · outbound

This paper cites Smartphones beat dram drum to meet performance demand, 2021.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Smartphones beat dram drum to meet performance demand, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.239712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.895668Z digest=sha256:53f13615e3cfc82a862cf839784d8a203152332bfa493f760e656e79e7660214

Observation c977e7a0-181c-496e-9c1d-4172fe11b670 · outbound

This paper cites How much ram should a phone have, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking How much ram should a phone have, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.220848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.901181Z digest=sha256:8a9e1ca2b5aafc9668f40c6cd029a2280591e003b4555af4b747086580a6a272

Observation 6f33b5aa-e1b5-4aa1-91f2-422074d7cdd6 · outbound

This paper cites N., Fan, A., Auli, M., and Grangier, D.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking N., Fan, A., Auli, M., and Grangier, D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.905750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.905750Z digest=sha256:fb404510aec32bbb46428aec3326c60eff62d42b3e323131ebe9481ee94b2612

Observation 54d12ab3-7d07-47d2-832b-b0e20c642ee6 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.910387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.910387Z digest=sha256:fa1081af79b447185372af07ea37c7621086ea1df588f09d42e4256f19fa1999

Observation aa89924d-8db7-4aa8-8592-bfca4782e7e7 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.915251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.915251Z digest=sha256:6bfc3a56f1ef305e8352fe4a91ac60714ff6c53faf8f6b06ec3b9730e5111a1f

Observation 53c840b7-2ac2-4992-bc00-60d1e7fde47f · outbound

This paper cites Code artifact for efficient llm inference using dynamic input pruning and cache-aware masking, March 2025.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Code artifact for efficient llm inference using dynamic input pruning and cache-aware masking, March 2025

Reference 11

Resolution
verified exact
doi, observed 2026-08-12T04:30:36.183513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.919920Z digest=sha256:262e96e07831e6be37870eb91fcfa4bb0b1987802fcc3b999db1f76c663e3a75

Observation c96977b5-55f0-4e86-adc9-c0a463f85676 · outbound

This paper cites and Alistarh, D.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Alistarh, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.924610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.924610Z digest=sha256:b19d4148584daf2042bfbd779d8a65fdce82e37f4abfa6a39f77bc6a073c6eca

Observation c5a7f179-82da-4fc0-bfa6-429a915aa293 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.929393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.929393Z digest=sha256:0c099a4a6c5bb1353b27f53e60889cbc92bd6689c2b9f3df0b3e1458c6d808e1

Observation 0c69a890-7a6f-436b-a286-0c27c9a20853 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A framework for few-shot language model evaluation, 07 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.934418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.934418Z digest=sha256:d0618b74afce9789106cacf3b3fb43e2dfd291a4624438c07d8640b4f6d49182

Observation 9a84b3fb-b85c-4189-829e-062bea9a86f5 · outbound

This paper cites W., and Keutzer, K.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking W., and Keutzer, K

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.938996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.938996Z digest=sha256:9df1265d1ed06ae765fd816011924da780fbd9eafc1affde4871cbadb060241a

Observation a52cae37-1f11-4309-9d1f-2f0678d8cf28 · outbound

This paper cites and Lorenz, J.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lorenz, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.161924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.943437Z digest=sha256:c4310fbfb4b361ed642e7f089efbdbc290be138412a0788a5b0674fe84af8fab

Observation de89ce20-02e1-44f0-b1f1-3b77bbc824a1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.948004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.948004Z digest=sha256:0b0cd634d363a13bd4c220fd22f6f43c19854c7dea456463b2d00d158becef78

Observation c159906b-805b-451f-8773-775215900693 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.952779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.952779Z digest=sha256:c7f4869bef4f08a5f19f8343d65267c3659edfed0a8aca0fa304e73ef7cccaac

Observation 980c6f8d-38d7-4218-89d8-32e703c6c764 · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.957853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.957853Z digest=sha256:0416e57285857033a4c6c83ca69c99ab5ae68652cda7da71e3d137e63b68170b

Observation 96bcb492-48e7-4f5d-9a59-c9ddd76ae092 · outbound

This paper cites and Lin, C.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lin, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.146110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.962685Z digest=sha256:8890d74ed76af9180b5bab54cc19c90ce86159b0549600fe22d8047d2fa45b0c

Observation 7b972b95-ad44-4721-912b-66af0df3758f · outbound

This paper cites Challenges and trends of sram-based computing-in-memory for ai edge devices.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Challenges and trends of sram-based computing-in-memory for ai edge devices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.130128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.968421Z digest=sha256:3e4b414aa76cece821fec7b2ab8a9146e2edce49134b9e9912a1900a24b35a60

Observation 1ffd6fdc-b3bc-4870-bc4f-035d3326d324 · outbound

This paper cites Mistral 7B.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.972943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.972943Z digest=sha256:3aa4ee61b44f3192b125680a42f0ea00d6e23fa860f4bd5f852207e95c514831

Observation f1b47f41-f004-43be-98c0-9698e15ca0fc · outbound

This paper cites Pruning vs quantization: Which is better? Advances in neural information processing systems, 36: 0 62414--62427, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: Which is better? Advances in neural information processing systems, 36: 0 62414--62427, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.113815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.977465Z digest=sha256:b9adcb2f94794e63daa1ecc0ad48da6b51e0d3ebcf9e28113cbe03e764b81a89

Observation f0f7afbc-d75c-482f-8ae6-ecfe48b82a42 · outbound

This paper cites Pruning vs quantization: which is better? Advances in neural information processing systems, 36, 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: which is better? Advances in neural information processing systems, 36, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.096028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.982024Z digest=sha256:907bc433b467ec30c11aebaca362846622d0c7c143ab41a8cad3cb0dd0cc9a14

Observation 347859e4-90cd-417d-9aaf-18d7d12244c1 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.987620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.987620Z digest=sha256:18c38cfd6fcef3f61c9a79657c8f6ea9236db5953de911b69b043c457a7230e3

Observation 6a1c6410-ba49-4e31-994f-f5ca55b5f7af · outbound

This paper cites Apple nvme vs ufs 4.0 storage, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple nvme vs ufs 4.0 storage, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.068082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.992271Z digest=sha256:f3cf4b177d9b5675472f0dd5ea609f4002193d694beaf29f96296bfcc3d20c52

Observation 5d5c0d83-edb6-4a88-b3f4-a6c401ad461e · outbound

This paper cites CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.996733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.996733Z digest=sha256:19b3a38962cd2d3e051e73082025ed440b5179f7f1da2d2d10826524a0fd8c04

Observation a047825c-5268-4438-aacf-4ff4eee851a3 · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:37.050115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.001626Z digest=sha256:dd97451a6c9554ca97ff1f841d0460c9d2d167f77f603ce9b15617e7dae6c6e3

Observation ea440abe-924e-4ad6-9b97-e15a61138d6f · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Deja vu: Contextual sparsity for efficient llms at inference time

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.006409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.006409Z digest=sha256:a203ad69bd968fbd48fd537216bba9d932f0d94e861a771ce710276e013b149c

Observation 40c17f66-d939-4db3-a385-2538d8109a1c · outbound

This paper cites and Vassilvitskii, S.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Vassilvitskii, S

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.010972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.010972Z digest=sha256:8ebf5a9d0e158b98066279a56f9aa1824681ed83185207ed3274f53de8068b3a

Observation 661d8d82-7e10-4393-8148-858695803850 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llm-pruner: On the structural pruning of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.015725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.015725Z digest=sha256:866bd4e1354acc2430f0ce586d9f72762486fcd4592885dd6aedc3c4c7bffdc1

Observation 33594af2-a8c9-4ee6-8f9b-88a2eb90954b · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.020982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.020982Z digest=sha256:45169dc0c539bbc1b21ad4407f21dfa5bd209f8573a815afb5b56745c77618c4

Observation 54c22a46-6e6f-4636-979a-6258aa59bcf4 · outbound

This paper cites A White Paper on Neural Network Quantization.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A White Paper on Neural Network Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.026374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.026374Z digest=sha256:c9cb077084da16b1dbe3c5ba0ae775fdc7398d95ec7c4ef3fa60dc5a07cfb030

Observation 77168a8e-e847-4c47-bc55-aa7753c693a5 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Comprehensive Overview of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.031099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.031099Z digest=sha256:986de75f7c0e4f16838da1f0f70706a4b23e727801e535b26f1fe5b4fd918f4b

Observation e87065d3-03dc-4e94-aa3c-bf575b49cace · outbound

This paper cites Llama 2: Early adopters’ utilization of meta’s new open-source pretrained model.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llama 2: Early adopters’ utilization of meta’s new open-source pretrained model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.001568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.036146Z digest=sha256:16616b6465a34ef9faf4ceaab22d0e57df05ba172048c6b1b6c8d97046b147e2

Observation 0808ae4c-4854-4cf7-815b-413bff1b881e · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:36.986475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.040734Z digest=sha256:f71ff8be5636afb7a5a8110b86c9cc7cee36281be85457de66a4ff3e5e118628

Observation b564d181-3870-4d71-91a6-bba295237c1c · outbound

This paper cites Effective mimicry of belady’s min policy.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Effective mimicry of belady’s min policy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.970840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.045822Z digest=sha256:1f339e1cc6a9ab341e307f4df4d90b0f09f5713f187e91a4f4314520077a4c0c

Observation a1f17811-df1d-40ef-b74a-c28de383297f · outbound

This paper cites R., Hestness, J., and Dey, N.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking R., Hestness, J., and Dey, N

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.955021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.050610Z digest=sha256:b36707ef4d8197628b78b8e98fb3c303af5e576072e06eb9f2765f259d754c91

Observation 9653c257-04d8-4d09-be4a-8622a7579b98 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.055200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.055200Z digest=sha256:f8691cca4c46bf20b33e6f64fe9ad0cc7157f64d3a323ff46dc637539d2bc7a2

Observation 0727e386-f3a9-497b-8cf2-561e39e4bada · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.060140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.060140Z digest=sha256:56a531e3425f6eb9409cb9e95bbd6c65d37e88b88a68319b5daee47dae6cdea0

Observation ba622ce4-2691-4519-87dd-712870577f6a · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.065313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.065313Z digest=sha256:370c3b482ccf7a51debb5f274f771b597f3fc39c13aab252e51cf62cdb2f4bf2

Observation bb59a358-0219-40a4-8d81-8de302534ebf · outbound

This paper cites Sparse language models with relu activations.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sparse language models with relu activations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.807249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.070513Z digest=sha256:e5bfd083b6fbc9c45f87e95936b0bf08a64852b2fd3c2183984d2e167899412a

Observation a0031b78-b421-4dd9-8280-2c322a9c32ae · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Simple and Effective Pruning Approach for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.075170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.075170Z digest=sha256:595dfd367656e32aa649b735fb7bb4b4ef2606ac119bfcc2925e815d9ba21d59

Observation 3088ac87-2ccb-4896-9e5c-2459bfbdc331 · outbound

This paper cites GPTVQ: The Blessing of Dimensionality for LLM Quantization.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTVQ: The Blessing of Dimensionality for LLM Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.080091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.080091Z digest=sha256:7254099275dfac093e46c5cfb840cb4a7cd7afc9ae3a242bc3c59c742a07dafb

Observation 6655a9d5-9afc-413a-8fd7-e7ce627e7381 · outbound

This paper cites The LLM Surgeon.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The LLM Surgeon

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.094555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.094555Z digest=sha256:7fb793c53be743fff6e05a03aebb2d84d0d5fc106fd8a0eea6d2393e51e4f63f

Observation 170d6ec3-2d9b-4302-848e-39d0cc1729f8 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.790571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.099571Z digest=sha256:43014717826e057a854f60eb31ed03a50094011c0e08c22dd3b476f4c2310230

Observation 9bc7e5cb-0308-46eb-8502-a7c4732ba1c3 · outbound

This paper cites Apple silicon --- Wikipedia , the free encyclopedia, 2024 a.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple silicon --- Wikipedia , the free encyclopedia, 2024 a

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.773992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.104470Z digest=sha256:764ff1bc50310da7a6da8190d97d8101dee6b82ae49e468420d59d5d8746ef8f

Observation 33a9ae10-e9ca-4ed8-98e0-17fd196aa542 · outbound

This paper cites List of qualcomm snapdragon systems on chips --- Wikipedia , the free encyclopedia, 2024 b.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking List of qualcomm snapdragon systems on chips --- Wikipedia , the free encyclopedia, 2024 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.754752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.109317Z digest=sha256:6503fff6b26afcef672a86f82b90388a414665ebc342b0d83a054b76501bb857

Observation 3318d1f0-5035-4b36-9cf9-243395fd3c4e · outbound

This paper cites Flash memory --- Wikipedia , the free encyclopedia, 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Flash memory --- Wikipedia , the free encyclopedia, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.738500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.114390Z digest=sha256:028f5e2ba090d022d209c58b67d0852e54a6b1e6c853ecef3c1aa658d93e2244

Observation 8c7d2330-15c8-4845-b7b3-f9906a84d1b9 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.119613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.119613Z digest=sha256:5c7d43c71239b34576cf2481a27930de5c1abd5495a483d47802256a4ca12fc1

Observation 08c025d0-cc7a-4893-b154-3e362a848f6d · outbound

This paper cites OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.124872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.124872Z digest=sha256:8b667bbfcba9de7d802b85b7a828c288ec1223af5b88b5b1401b3a2187b87195

Observation 31f10859-551d-4b20-81e3-9e7caedddce9 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OPT: Open Pre-trained Transformer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.130490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.130490Z digest=sha256:7dcae7248d4fefcec8631cd6427670f3e06ab9f58eebc92aa798f55031c346f5

Observation 667f0a38-94ac-477e-93b1-340159c08a00 · outbound

This paper cites A Survey of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Survey of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.135498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.135498Z digest=sha256:58fc6340540b89f19b608dcd59d4b6c627aba4b9adcf4a289b90b6300745c887

Observation e3d4b92d-51ca-4b6d-b01c-ea9de18e957e · outbound

This paper cites write newline.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.140273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.140273Z digest=sha256:873cfd133378d785de3fa34fe34df923b19971a9e9d930db4fc321e2dcd02d55

Pith citing papers

Observation 57e8b28b-7346-4cbf-9aca-f48e6a169591 · inbound

RAP: Runtime Adaptive Pruning for LLM Inference cites this paper.

RAP: Runtime Adaptive Pruning for LLM Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:21:35.629058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T13:20:41.739571Z digest=sha256:f1f8f82cc93f7b8f426e9e17dca6554441bd83c01934f5adb4e468aa3437a5e2

Observation 244a3fa0-3f99-483d-b108-e6139357ad16 · inbound

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures cites this paper.

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:49.289390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:49.289390Z digest=sha256:8f723a9df4f671631b7f98f9a9b700cc89da2f4dfbab4a326bc02567ef5c91a5

Observation b1490825-f9dd-4591-a84c-a20a207d8aab · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:56:59.496372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:56:59.496372Z digest=sha256:173e96d052bb9d5dfcf376bc2b39ce98447aa52ff204702f93d863f7c188d5b9