Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Language Models with Zero-Shot Adjustable Acceleration

As of 12 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2509.01190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01190 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:52:44.664299Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfa35c9f-fdef-4003-9f08-d30a3be4ddf8 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.609841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.609841Z digest=sha256:be88d4d18954f4f02131d6eaec8fce6fc13985892cc552e0a412a2abfc1b88b7

Observation 9617c512-fa86-40b4-ac56-e644275340ea · outbound

This paper cites SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:52:44.760293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.629911Z digest=sha256:7cdb1311c9569e62025cdc82fbf5d6da7f7dca810aa2ad7a3a6062ee06eb0084

Observation fe5ab983-814a-450f-90b0-d9eca33518af · outbound

This paper cites Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:45.011281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.636353Z digest=sha256:fb7e8508806658ecec588fce05266cd99b78da982c7c70ff10f2391effd221b9

Observation d5bd6782-a633-4b16-9a0c-2db2ab7e5628 · outbound

This paper cites RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:52:44.721825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.641875Z digest=sha256:b76be06e415e4b21f2f48ccc66db9028ecc458eb6dccfa61ecfcacbe7b999935

Observation 3e917123-a917-467e-bf33-8a4af29a9b71 · outbound

This paper cites InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.654033Z digest=sha256:c066f17dc136213a03e89bcaaa517830c67a03260da3a899f0c304ac9ae18900

Observation cf8f6fb1-49bb-4c5b-b2fd-b5ac5bbf8d0d · outbound

This paper cites In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.938126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.658328Z digest=sha256:bf0e0848a3f5fc5a1cecb9724983081b7840e083255411ee27e7e6f8d5a35e03

Observation 107c0efa-055e-4229-8830-5e8482b3c100 · outbound

This paper cites Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.917986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.664299Z digest=sha256:cd1e744c168632c6d2e5b972bf9943e798fdefb253b53066fbf04f14afc4b9fc

Observation c16095b2-c899-494f-9a47-984c45f04b98 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.604653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.604653Z digest=sha256:32d320f6953f6a21f3e1d5ee0b8208739ccbbaba2446916b1764370623b5fbdd

Observation aecc832f-0f19-45a3-b2b6-e1f0bdcfac3d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.623223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.623223Z digest=sha256:794b11dd3b313681a59748753044d846fde1f7f5c9872c7ac367ac34dd08d9f6

Observation f0b1ab8e-093c-472a-a808-1026bf03a6c1 · outbound

This paper cites InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.992848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T12:52:44.649209Z digest=sha256:17f761943920d0062b3d60b73b61ed56dafbd8f62a378d51cd3a626897c8ea9f

Observation 379ad255-dff5-4788-8671-cd5f7f99c0d7 · outbound

This paper cites Token Merging: Your ViT But Faster.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Token Merging: Your ViT But Faster

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.591356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.591356Z digest=sha256:ff7ea6f8fbecf55a7b11f45099ac02ad68dfa52514f825b907715b6478c0fa02

Observation 1eed93bd-164e-4734-afbc-880163b626f3 · outbound

This paper cites LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.617762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.617762Z digest=sha256:30806d3fdb9d3367957dc4008719bb361ae0590c025ecaa4da733cc85341b1d2

Observation 8bfc1734-8aa2-4359-a406-3592205cbe19 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.599095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.599095Z digest=sha256:84441f93f3f320211724be36cc8c7a3204d588e4c5708e2d4ba64ad76b552780

Pith citing papers

No inbound Pith citation observations are available.