Pith. sign in

Paper Citation Record · LEDGER

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2506.09104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09104 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:03.961067Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T23:28:12.790404Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.258102Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21764a78-b4eb-47e1-b8d1-ce16dd3bed6c · outbound

This paper cites Paretoq: Scaling laws in extremely low-bit llm quantization, 2025a.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Paretoq: Scaling laws in extremely low-bit llm quantization, 2025a

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.768971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.768971Z digest=sha256:ae41e440e1a0de3e1e2beee3330cc43efdbbef984dd504a937f696c3fec43379

Observation 8c4cb3a6-46e1-4183-bac2-86cdaabd1031 · outbound

This paper cites Evaluating Quantized Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Evaluating Quantized Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.889275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.889275Z digest=sha256:6e5d89e858a3206727bcd3b0e8e4809395d9423a474bc3672ac16cf023807590

Observation 87080d35-f0c3-421d-aa94-a84a33b5ce14 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.029383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.029383Z digest=sha256:85d9fa9d91ab0980e4fc28f9a0b9be3b70a16d3acb81c8cfa2adaf8771525663

Observation 8e479b64-4769-4745-8936-23e0b4357318 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.157167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.157167Z digest=sha256:6e0ffb256ee6f3f848ed6e397cbb07c7a5edc4d21c3818969b003138787ec0ce

Observation fd29aa23-4d4d-4923-ac68-3e36efdde6b6 · outbound

This paper cites BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.290534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.290534Z digest=sha256:1f618e6f07354d0e4bf30d5e443ca3718f2a16caf597bfe1265a9182ad0a6fdf

Observation ea7a7258-9f9b-4821-b58d-3b2a190050b8 · outbound

This paper cites Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.353964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.353964Z digest=sha256:c118f498f8b371aec7947a5a5940794220aa97785746b1d25c3564b11f9f32f1

Observation 3a60f12e-39ee-4ef5-8bf1-2a97c447cfb6 · outbound

This paper cites The Llama 3 Herd of Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.408635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.408635Z digest=sha256:0842e0626a281beded51e4874b8f0f29f22344f34262027b69328b7eb99456cb

Observation d16d2e05-dc17-4e7e-afc3-1719d2005b6e · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs DataComp-LM: In search of the next generation of training sets for language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.563789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.563789Z digest=sha256:f55c368bfe47c124083755084c704cef4346930744040a3737c7cdfa62d9a8b8

Observation 07f79d10-4c5c-4fd6-9026-00fa5c75fbbf · outbound

This paper cites Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.737130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.737130Z digest=sha256:889799685f66eb45609cc1c9383583140438cf7fc94e2a28617b44e9160c495c

Observation 4c4e4d04-5021-47de-b862-c01e66c963de · outbound

This paper cites A Practical Mixed Precision Algorithm for Post-Training Quantization.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs A Practical Mixed Precision Algorithm for Post-Training Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.801268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.801268Z digest=sha256:547cd81a16869ee042c7076d9046c78009f3693a78f8615b32b1af5571d49caf

Observation 35e5ee0c-d156-4529-be51-ea568a4697ae · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.858593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.858593Z digest=sha256:389a6f26a6cec68631c519c56dd1e380886096a7b50d8871101eb24d55bdb15b

Observation fea0784d-eb2d-47e7-b56c-067165aa9c26 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.905843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.905843Z digest=sha256:3231fafd89542bcc27f688d080f49512fcb5c9b474bce09b3f7cd751f96e1fa9

Observation fb904247-fbb2-49d0-99c1-9eaac1453974 · outbound

This paper cites an unresolved cited work.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:04:04.513967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:04:03.961067Z digest=sha256:b634f04472c64c1a2b7ce02034942d54006177c096499f05df891553cc7f0560

Observation e30bb9b0-cc2f-4e27-b647-e26ed7790c2b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.637756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.637756Z digest=sha256:95718cc881917007c957a2b9eb2ef09f2264002fd60f0a2cd7b2d221fe943c6e

Observation 9587349f-bf95-48ff-98dc-0462929cb415 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.689898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.689898Z digest=sha256:f6f320e14d2b840edeb541c9744ad92e5797f43c101d00fb428fd3777c2df463

Observation e385ac98-30e5-4afb-b474-7c28198c8c49 · outbound

This paper cites doi: 10.18653/v1/2020.emnlp-main.494.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs doi: 10.18653/v1/2020.emnlp-main.494

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.239711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.239711Z digest=sha256:75a85bc00a0d187a52f0678a7300cf32ecaa042404804d4f5795cb4831da5487

Observation c180ca09-9f15-48a2-b7ec-15c3a2e27e4e · outbound

This paper cites Sharpness-aware Quantization for Deep Neural Networks.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Sharpness-aware Quantization for Deep Neural Networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.956190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.956190Z digest=sha256:f0dfab4ef1aa4e54264a98881f8d29ab51e21bed61a4c71d15bf6e0d90eaa21e

Observation 05946ffe-d495-459e-abf9-48c568fa523b · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.811644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.811644Z digest=sha256:6702db1461841f57ed6594b311955bcce7f650c4521d13a58ecfd6df4e0a08f6

Observation ab2c83ec-7cc4-4bd6-b8aa-0fcee475aed7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.098057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.098057Z digest=sha256:f50ede14076bff4741ee31da76f9e3c3d73a88528148e56d8d7b6f93696207e2

Observation 6b5bce84-5d51-42b6-9aa0-c786512bb809 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.467515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.467515Z digest=sha256:50c45fda42094770f1a5b77b10dc5ba12926f5c1fc36d4178eeaadb84bdd482b

Pith citing papers

Observation a37eb065-df34-407a-ad1f-1c5b1d403cc9 · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

Reference 148

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:d8105cc08c869bf495ecec46090aa17bfb99278248eac04c40d271f43bd2e3c7

Observation f89623b7-ee66-4c66-9230-e0a86ebf87bd · inbound

Apertus LLM Family Expansion via Distillation and Quantization cites this paper.

Apertus LLM Family Expansion via Distillation and Quantization Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.259669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:27:47.476242Z digest=sha256:55ec29ecb8175e6a9cd03bce184402ff2be4898ac0fb3cc7336321084b3ec728