Pith. sign in

Paper Citation Record · LEDGER

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.12216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12216 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:32.136821Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38e2a364-1366-44e6-9eb5-38f3223ad163 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.996022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.996022Z digest=sha256:0a4905ad12effd5880ec06b8a1625649479bfda70650fe8ceca029e63db1f590

Observation 9ce29c8e-504f-47db-af1b-a2b46962e63b · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.000215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.000215Z digest=sha256:4fc8261b23ac08168c1ef1564275e7b8ca09d696cb164491d5d4814c254a86ae

Observation 7a2efdac-6b4f-4fd4-a41a-6e89c241eca1 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.106556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.106556Z digest=sha256:87b6ceea9d841ef28f326a6cbffc9e71d5b03a692d8cee69cc40473bbe68a6f6

Observation 4b9f4fd5-1703-47c2-8ff1-acd17001c7a7 · outbound

This paper cites EvoPress: Accurate Dynamic Model Compression via Evolutionary Search.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models EvoPress: Accurate Dynamic Model Compression via Evolutionary Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.115409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.115409Z digest=sha256:7e7bebd0d1f26687a32102d3ae63a62a38584e4460e0d4a09cb5af42e682bfdb

Observation d6712259-3d05-4354-8759-50afda2488a3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.123193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.123193Z digest=sha256:ff4b71b6ec5c6efc295d3d50a6a59c646cd3904268c608b79103de743fd5cef5

Observation 0ac7c410-f328-4c0c-b2d5-37322410a269 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.127470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.127470Z digest=sha256:652f7e2e0af1ba60c96bf0666a0f40db37fa17ea7a9a002443df8c67ceeeaf35

Observation 5556c198-136c-4e34-ac83-1475d5b3fda7 · outbound

This paper cites Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.280711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.132173Z digest=sha256:f4462304724b523d2405af98fb12d9ebc1712fffc628b01cfaf4241a96e92a63

Observation b9fb8bea-e17b-4f0e-a736-526f62c90390 · outbound

This paper cites SparseGPT generates spar- sified blocks with varying sparsity levels across layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models SparseGPT generates spar- sified blocks with varying sparsity levels across layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.268129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.136821Z digest=sha256:5cd1f5fe5c3afe044e57bc35cb900f74e5d92bc2a84e7b9c4e9dfd1eb11000ce

Observation e7ef7af7-77ce-4ae9-8a5c-52ce8f7785a0 · outbound

This paper cites In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.292280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:32.119355Z digest=sha256:28d6fafc6a57f0a2ba2d0d967b37bb74753966670d1ee7c7bdf2104cbecb8cc2

Observation 3a0d2829-fc77-4fce-a0fd-2f7d1ddc7fc8 · outbound

This paper cites Pointer Sentinel Mixture Models.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Pointer Sentinel Mixture Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.101934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.101934Z digest=sha256:e5aa5e38c07b1747b004fb78b3b142fc45a1f81178cc2d58bb9f963fe0398989

Observation 4bdc28a0-2ff6-4cb9-b7f0-fcd5f2a05e0c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.986746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.986746Z digest=sha256:5288d77e63a930b904c9f49d298ac2b7a6820ee230c68635b3432071d6f928cc

Observation 33b7391d-9ae5-4c39-82ad-7f837963e7c7 · outbound

This paper cites Weight subcloning: direct initialization of transformers using larger pretrained ones.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Weight subcloning: direct initialization of transformers using larger pretrained ones

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.111122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.111122Z digest=sha256:92ac6f14e4b738696a8ca11f4a7e512664a6adac568abd997d0ae0f401dbe025

Observation 45bf08ae-5f02-4f3c-a600-d612d4339a6c · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.991534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.991534Z digest=sha256:32c12fe820a9cfe1e5e3896e93fb6d8ff34f0af7efc95bc57fdf8f6ddcb03fb9

Pith citing papers

No inbound Pith citation observations are available.