Pith. sign in

Paper Citation Record · LEDGER

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.13735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13735 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:55:41.247033Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64528687-9085-466a-988c-25f3391937f9 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:39.877162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:39.877162Z digest=sha256:c70f477009e3006f1d86c12c711c230b342a889e6a0a03d85faf5c15ec2b4dc2

Observation c93a4183-b56e-4ef5-ab2f-0cba141bdf49 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.314783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.314783Z digest=sha256:8dfdcb9f7bd80b14543876222209e498c350b53de4dbd8f74a4d904e75f16851

Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.433775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.433775Z digest=sha256:78e7fba3990310f2a408006e5032e1aac7df22810bd67c1fbd3de05c777a976c

Observation db20af93-eb7e-4b9b-8bb2-9d879a52b75d · outbound

This paper cites Norm Tweaking: High-performance Low-bit Quantization of Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Norm Tweaking: High-performance Low-bit Quantization of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.548166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.548166Z digest=sha256:fbe3dfb5bf44feefab6d450c84f71e7f8881f04d2db709e1b5f24685c0e99186

Observation 19cb1add-ad6b-41f6-ae63-a819d29e2108 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.665371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.665371Z digest=sha256:81a454c865f0b4edaa0b07f47866e3b2cf496f87f601fc89979cf678dad438de

Observation eda8e874-d59f-4343-a954-423ffee4a1c1 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.779355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.779355Z digest=sha256:8064114f4bbe9020fc3a99b5b0d3a6e2c22329cc3731cc53f5432aef065ff5ad

Observation 74f57f1b-a497-4f8b-a34a-3c9c416713ae · outbound

This paper cites Distilling reasoning capabilities into smaller language models.Findings of the Association for Computational Linguistics: ACL 2023,.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Distilling reasoning capabilities into smaller language models.Findings of the Association for Computational Linguistics: ACL 2023,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.897846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.897846Z digest=sha256:3110ab4c84229f0f39409ab2e270d12b480624a80761bfe3bcd16019db6df881

Observation 641e44b7-2516-4e54-9675-b172ff832c84 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems A Simple and Effective Pruning Approach for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.015173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.015173Z digest=sha256:4e38b07a6c86032e0b64330d402c48e0cb43efc49d15a5d8ba81b98b3cd6e80d

Observation 66defc3c-1c96-409c-be9e-3130e323c914 · outbound

This paper cites ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.125011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.125011Z digest=sha256:c1600a5f74ad99822a6843882879a15558022fb2a70cfc4b96e710cb54468821

Observation 72c9aa9b-2a55-434a-8003-4f91bdf4d635 · outbound

This paper cites LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.247033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.247033Z digest=sha256:d26dc72f85e627ad52c556fb35c9992a5b01cdddb396bae9352df91826c1b114

Observation 28dd328d-0752-4925-a4a8-742b323b8eca · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.168571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.168571Z digest=sha256:96e90b836557746d286d7dd05eb7773d81f48ff73a9cb786e11e9acafa9f4877

Observation c3983ad0-d2f4-4b19-9723-3a2430f34d74 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.066366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.066366Z digest=sha256:db39ee1edf4360b81e8e10d5c175eb4a84c5718fc28ddad061cc9b26ec3adf13

Observation b19c3f69-c03e-4e56-aa82-9f815b64f46d · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:39.951387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:39.951387Z digest=sha256:9500bd74e2f10ada2333844322f11e5a85f75fe6949a5d46b128645cfa0e2b05

Pith citing papers

No inbound Pith citation observations are available.