Pith. sign in

Paper Citation Record · LEDGER

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems

As of 21 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.13735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13735 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:55:41.247033Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64528687-9085-466a-988c-25f3391937f9 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:39.877162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:39.877162Z digest=sha256:c91043a5de7ec8c8bbc9aa3fc79bc708e68685572a76b003e48233ba8c89c35f

Observation c93a4183-b56e-4ef5-ab2f-0cba141bdf49 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.314783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.314783Z digest=sha256:d4bfe3cb6fd2a18ef7fa119ed4054306e11f2bd696cafe1e8874f54820df5870

Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.433775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.433775Z digest=sha256:1379776afd53633a09eca3bce5b350fa2853b7d4fc482f9596286b156a9fc4d3

Observation db20af93-eb7e-4b9b-8bb2-9d879a52b75d · outbound

This paper cites Norm Tweaking: High-performance Low-bit Quantization of Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Norm Tweaking: High-performance Low-bit Quantization of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.548166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.548166Z digest=sha256:d06bb47baaf559fb5fd17f4351ba7ab498746d759ad4b0c0ab7654deff830054

Observation 19cb1add-ad6b-41f6-ae63-a819d29e2108 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.665371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.665371Z digest=sha256:f67241ea5f28f6984fe865da087093e067e572404269e264b472565f603dedc2

Observation eda8e874-d59f-4343-a954-423ffee4a1c1 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.779355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.779355Z digest=sha256:3ab8f4c0d98ebf574bdd302b8650a464734a7e27fd236de436cb30dd9c483657

Observation 74f57f1b-a497-4f8b-a34a-3c9c416713ae · outbound

This paper cites Distilling reasoning capabilities into smaller language models.Findings of the Association for Computational Linguistics: ACL 2023,.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems Distilling reasoning capabilities into smaller language models.Findings of the Association for Computational Linguistics: ACL 2023,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.897846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.897846Z digest=sha256:083a3adb1e2bde3ba4db26e16dcc94f06231be4f8a475d4fda666d7fad617048

Observation 641e44b7-2516-4e54-9675-b172ff832c84 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems A Simple and Effective Pruning Approach for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.015173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.015173Z digest=sha256:34a44935c49fd0cd962d1381840820f93661ffa4a16ee0a100bfa52278e31599

Observation 66defc3c-1c96-409c-be9e-3130e323c914 · outbound

This paper cites ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.125011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.125011Z digest=sha256:5cf347deaea84d8797a5b0fbb24edb4f5c20ba8219c8006a79579d892e944505

Observation 72c9aa9b-2a55-434a-8003-4f91bdf4d635 · outbound

This paper cites LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.247033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.247033Z digest=sha256:e3b58e0ac93c821239ab0d1c204d7785f1cae5c7a8ace1a85eb7e707916bd0a0

Observation 28dd328d-0752-4925-a4a8-742b323b8eca · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.168571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.168571Z digest=sha256:4b377eb862480dff96f6890c0917c26b350eb394e45d159fa7446fb64946054e

Observation c3983ad0-d2f4-4b19-9723-3a2430f34d74 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.066366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.066366Z digest=sha256:f5c530851de17f71c0dbca0f222c429fdf021285b0146d13e163785f2cddb77e

Observation b19c3f69-c03e-4e56-aa82-9f815b64f46d · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:39.951387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:39.951387Z digest=sha256:30c7f4f3108a6c335875ce3e7113ef3669a4d75dd805ae2fdb70cb7c55578f2c

Pith citing papers

No inbound Pith citation observations are available.