Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Language Models with Zero-Shot Adjustable Acceleration

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2509.01190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01190 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:52:44.664299Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfa35c9f-fdef-4003-9f08-d30a3be4ddf8 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.609841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.609841Z digest=sha256:e65de992e0038472db87955fbfad6674943ec5310159a97c9be1301e8ac40ca5

Observation 9617c512-fa86-40b4-ac56-e644275340ea · outbound

This paper cites SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:52:44.760293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.629911Z digest=sha256:a67e024647dd34fa81ca1b3d32a5c3592838c72cd9f7020f7a147f8aaaf29e50

Observation fe5ab983-814a-450f-90b0-d9eca33518af · outbound

This paper cites Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:45.011281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.636353Z digest=sha256:fd173b228795b121ad564628aab1d85560700d8b102a43b59424ae189807f24b

Observation d5bd6782-a633-4b16-9a0c-2db2ab7e5628 · outbound

This paper cites RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:52:44.721825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.641875Z digest=sha256:a8369d7f314a4f0eb6ca292b0d2643cca9c298724851daaca418665ee658f511

Observation 3e917123-a917-467e-bf33-8a4af29a9b71 · outbound

This paper cites InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.654033Z digest=sha256:c8438b2cccb3146a639c995d4fd36d2e0474335a822901b3d0bbb7a2379bcc62

Observation cf8f6fb1-49bb-4c5b-b2fd-b5ac5bbf8d0d · outbound

This paper cites In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.938126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.658328Z digest=sha256:a21a2846bd5fb6060a809aacf6554daaf12b5fd7764028473b6cf42a44befe6f

Observation 107c0efa-055e-4229-8830-5e8482b3c100 · outbound

This paper cites Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.917986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.664299Z digest=sha256:3982c418c5b540572b9d6ce88c07cda75911250e92fb3d2473414b419e9bcbd3

Observation c16095b2-c899-494f-9a47-984c45f04b98 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.604653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.604653Z digest=sha256:617b4749aedded81c89fc8d368be7c772648bbd74171fdf0a97b4a05fda8b425

Observation aecc832f-0f19-45a3-b2b6-e1f0bdcfac3d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.623223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.623223Z digest=sha256:5b1ed5611a0e26a3df0b3673dd3c706f0fba94f3ce4d9d7c2ab72784d96522ba

Observation f0b1ab8e-093c-472a-a808-1026bf03a6c1 · outbound

This paper cites InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.992848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T12:52:44.649209Z digest=sha256:204c03b49c8e8b34b048d537e8dae08de4461d523e33e68031649d4c1267546b

Observation 379ad255-dff5-4788-8671-cd5f7f99c0d7 · outbound

This paper cites Token Merging: Your ViT But Faster.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Token Merging: Your ViT But Faster

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.591356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.591356Z digest=sha256:914fc46e57ab3c4d8be2a3e1e0bf32f14d94b41ada5c3578e32d641386727b4c

Observation 1eed93bd-164e-4734-afbc-880163b626f3 · outbound

This paper cites LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.617762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.617762Z digest=sha256:36c9347ffea04bac6c5566edf86bc582683df8b558c367eec394d8d1cb1bea4e

Observation 8bfc1734-8aa2-4359-a406-3592205cbe19 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.599095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.599095Z digest=sha256:65fb8e9efb165533437123b1b6324c08a157d8ecd25dc88c4a9d443c973063f6

Pith citing papers

No inbound Pith citation observations are available.