Pith. sign in

Paper Citation Record · LEDGER

Model Compression and Efficient Inference for Large Language Models: A Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2402.09748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09748 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:47:33.625921Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:56:27.554566Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d24ca4ff-f888-49d9-811d-5557fb4027a4 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:21:39.769822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:a9e4d96d042830186c8b6d609e49184ad6b0ca5e01dbba5ab8d5a45d7bb9098c

Observation 835075e9-1ccd-4deb-ae0f-2736f3b201c6 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.448971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:73aa20870ed96d735df395160867b065655b89777df8f1fe8ef11ac1244e14ba

Observation e42e9059-5993-418e-b343-be9a7a8d1a32 · inbound

SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs cites this paper.

SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:47:33.625921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:47:33.625921Z digest=sha256:c97a7a379c21352154c5b6ba415204a2c7156141b824c7a6444b7b6f4c8da057

Observation 672c3344-f7a9-49ba-95cb-9c901ddbcfe5 · inbound

Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones cites this paper.

Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:23.498590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:23.498590Z digest=sha256:5b480875dc0fa0c5fa9481c79c7c5e4c6e1407d85b2eeaf9ebbea4bf80e4ef08

Observation d0c3594e-3379-4cdf-940e-f6ebc5f120d6 · inbound

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code cites this paper.

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:01:55.806760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:01:42.228190Z digest=sha256:40d958b4332f774177123d06627f1623900038560a3edab20db9875125b3af14

Observation 0b1e59f2-78c2-4de4-aa9b-8cc03e553711 · inbound

Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training cites this paper.

Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:35.117902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:35.117902Z digest=sha256:d23817563814642140b5b308a3d8347dcc95fa425aa609897820293887f81547

Observation 5a70f26f-42a2-4858-8760-ff0151d0109d · inbound

Less LLM, More Documents: Searching for Improved RAG cites this paper.

Less LLM, More Documents: Searching for Improved RAG Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:16:19.195844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:13:01.397197Z digest=sha256:b6d7320c8d4b881d1fa7cadb654561f5cb615a5c3dadcb089cb226a3d6f88509

Observation a3eccf1f-c30e-4128-a5f4-2a3659e5b6db · inbound

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models cites this paper.

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:35:50.708362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:35:50.708362Z digest=sha256:49d06429db400f4e799141ad2b58ec122b035114d1256abfbd5b131d0ca29151

Observation 595d78a1-daf7-4559-a298-6c6e163dae39 · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning Model Compression and Efficient Inference for Large Language Models: A Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.557198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T11:27:02.902720Z digest=sha256:87a1712dea9b6428fccf8e6c56e7a37d17dce8c4de2486034c775da543086aee