Pith. sign in

Paper Citation Record · LEDGER

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.12216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12216 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:32.136821Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38e2a364-1366-44e6-9eb5-38f3223ad163 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.996022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.996022Z digest=sha256:45c194383410c371518d0ff7177c36e114cf44902f8a909f9d0c742ae82c1e42

Observation 9ce29c8e-504f-47db-af1b-a2b46962e63b · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.000215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.000215Z digest=sha256:edec5e959bccab46033f7ddbf2df09a80471a1c68ff02e63d13ab51789944133

Observation 7a2efdac-6b4f-4fd4-a41a-6e89c241eca1 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.106556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.106556Z digest=sha256:b91d2db556dd9c320a73c051aabf1e166b0896a101e24540307da19f47849fa1

Observation 4b9f4fd5-1703-47c2-8ff1-acd17001c7a7 · outbound

This paper cites EvoPress: Accurate Dynamic Model Compression via Evolutionary Search.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models EvoPress: Accurate Dynamic Model Compression via Evolutionary Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.115409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.115409Z digest=sha256:1b0928ba61213a32e3b0f8ce3d42e8a6bb85e1992775c595db76e3a27ecbcb39

Observation d6712259-3d05-4354-8759-50afda2488a3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.123193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.123193Z digest=sha256:a38d7a51ea1eef79e353f628fc21339dd3fc7bbffb228514e6194f54fdf8f70b

Observation 0ac7c410-f328-4c0c-b2d5-37322410a269 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.127470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.127470Z digest=sha256:be2e9976dd7d682c962609f71563f547b70026a12fb86c7982569cfdabfae926

Observation 5556c198-136c-4e34-ac83-1475d5b3fda7 · outbound

This paper cites Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Transactions of the Associa- tion for Computational Linguistics, 12:1556–1577

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.280711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:44:32.132173Z digest=sha256:f543ee97948200d8ba2d31b140e8cfef9cde12831f6bf74b5653f14a7cf38127

Observation b9fb8bea-e17b-4f0e-a736-526f62c90390 · outbound

This paper cites SparseGPT generates spar- sified blocks with varying sparsity levels across layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models SparseGPT generates spar- sified blocks with varying sparsity levels across layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.268129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:44:32.136821Z digest=sha256:a18cc8f3b0de529fc4fb699b19dd3cb869502b488256097cf22c53cba625d60c

Observation e7ef7af7-77ce-4ae9-8a5c-52ce8f7785a0 · outbound

This paper cites In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models In 15th In- ternational Conference on Scientific and Statistical Database Management, 2003., pages 141–150

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:32.292280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:44:32.119355Z digest=sha256:f9d6886e7ecbe8f6058c8cd36ad0502464ae39d81f62ae165ca80f25a01f304f

Observation 3a0d2829-fc77-4fce-a0fd-2f7d1ddc7fc8 · outbound

This paper cites Pointer Sentinel Mixture Models.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Pointer Sentinel Mixture Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.101934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.101934Z digest=sha256:4dd704466e7e032002ef9380c4c81ec7d954672cc891c83b4015b34fa4637913

Observation 4bdc28a0-2ff6-4cb9-b7f0-fcd5f2a05e0c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.986746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.986746Z digest=sha256:e50f5e68e23e5b536d0f8c622cd1f101fd83937b3075ac4b10d88a63cf0fd6de

Observation 33b7391d-9ae5-4c39-82ad-7f837963e7c7 · outbound

This paper cites Weight subcloning: direct initialization of transformers using larger pretrained ones.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models Weight subcloning: direct initialization of transformers using larger pretrained ones

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.111122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.111122Z digest=sha256:35ee906af04f33a582da4a4c442b3b91cbf6c74936ff058d00f1cc6e307907ff

Observation 45bf08ae-5f02-4f3c-a600-d612d4339a6c · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:31.991534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:31.991534Z digest=sha256:fa37f5493af03294e5abc32db80c4cf90586e12425953aee0362f684e3a82343

Pith citing papers

No inbound Pith citation observations are available.