Pith. sign in

Paper Citation Record · LEDGER

BAQ: Efficient Bit Allocation Quantization for Large Language Models

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.05664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05664 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:55.423637Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-05T05:41:13.869451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T05:50:43.590536Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d542fac4-a47a-4dfa-93b8-fb721f8ba1f6 · outbound

This paper cites Introducing ChatGPT.OpenAI Blog.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Introducing ChatGPT.OpenAI Blog

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.981600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:51.125122Z digest=sha256:42ca148f204593880d1b95a43322d158ba432c8a9794df1d7ff8334c9d34f8a7

Observation 01ae5068-8aa3-42a2-bb3f-e8dc07648b74 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.Advances in neural information processing systems, 36:10088– 10115, 2023.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Qlora: Efficient finetuning of quantized llms.Advances in neural information processing systems, 36:10088– 10115, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.215053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.215053Z digest=sha256:120296bc8d03d5137229af8b8afefb5db55c1ac0b111d72eeb18037972d528e7

Observation d37587dc-eb2b-445b-8843-73e5bf796cd0 · outbound

This paper cites OPTQ: Accurate quantization for generative pre-trained transformers.

BAQ: Efficient Bit Allocation Quantization for Large Language Models OPTQ: Accurate quantization for generative pre-trained transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.356767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.356767Z digest=sha256:3f0efc1e7d267e62249dd6bc4dedeb0412582333cb11c0453f4fc022a7e37d12

Observation c4638107-943c-4efa-842d-ece7443e4e08 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.557589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.557589Z digest=sha256:a81423dda5a3b1d8b26e271098ad65cb2744e83f6b9e2d24d806636b84266eda

Observation adbbb647-40a9-4e9f-b99c-bc4e58e2d344 · outbound

This paper cites QuIP: 2-bit quantization of large language models with guarantees.

BAQ: Efficient Bit Allocation Quantization for Large Language Models QuIP: 2-bit quantization of large language models with guarantees

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.643512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:51.709490Z digest=sha256:a37fd77d8466175519e7ac9348ef050ba231de04d91e9053c94a7b80014e6af9

Observation 56f249c7-4c10-407a-bcfc-4a6227bbbff0 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.833564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.833564Z digest=sha256:2df389182ff2b66a607bdaebe7c12cbadb5f449fe68c02727919134638406a50

Observation 825ef0da-5999-4b78-9e0e-5247cf497052 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quan- tization for large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Omniquant: Omnidirectionally calibrated quan- tization for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.035060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.035060Z digest=sha256:8635c2f9c59c54c297f1adfb3bdb0cc577a393ebb8a68ae683aed9fcebe6d507

Observation 991ba312-d4ad-4764-8e7e-2ebcca51e8a8 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.178887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.178887Z digest=sha256:299bc0ac2e265db33ac03ee18e6f525ca2843abfd3497ee04ad3ea745216dbc3

Observation f28f6709-fb4c-4d74-b87c-bdbdfe264093 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.315397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.315397Z digest=sha256:68a56379df95e5bc1a15eb21d5ad0d7202d505e0411bd149dd6003b3191fd729

Observation c08e9e6a-75e3-4440-a082-d8b03450aba9 · outbound

This paper cites Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.490197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.490197Z digest=sha256:7a6df9921a11e5c45b5c5ec5fd57390835ae7add54801fdc27d195569ba37996

Observation 35aa94be-e2a4-4b81-aa83-0b824d01744a · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

BAQ: Efficient Bit Allocation Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.624716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.624716Z digest=sha256:fc2cb20ed23fabbf0ecb507297a777fe462f4f56c230a0fd5c511933bed501b1

Observation 6b7e28c6-9be5-4042-b20e-ccb2a8fb9961 · outbound

This paper cites Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.749496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.749496Z digest=sha256:1663ce12b0bb88551bd3b5845b635c6929157e6f4f367d7ecb8e786543f82d39

Observation 52fbcbee-8dc5-48c3-b17d-615644050114 · outbound

This paper cites Optimal brain surgeon and general network pruning.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain surgeon and general network pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.376113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:52.920639Z digest=sha256:66b9052a58cc620bd42b8fa018dee5eb987707a7e0963bdc9aee5e263f46d1ed

Observation c011a7f5-878c-40ab-836e-7053d7aa0c89 · outbound

This paper cites Optimal brain damage.Advances in neural information processing systems, 2, 1989.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain damage.Advances in neural information processing systems, 2, 1989

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.069213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.069213Z digest=sha256:1ba322206df32e2ef2b6e11fc363f01f895924e8095631c69ed2f02254996788

Observation e00c5e90-9f8a-4caf-8480-2c0e09f4f89c · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.061585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:53.189419Z digest=sha256:3395e7bb47fb7af951facffb0e74d35e7a874ddfcb5e168a893394312a7cfcf7

Observation 20cf8c7a-20cf-4c73-a47d-3e33a7b449c2 · outbound

This paper cites Optimal brain compression: A framework for accurate post- training quantization and pruning.Advances in Neural Information Processing Systems, 35:4475– 4488, 2022.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain compression: A framework for accurate post- training quantization and pruning.Advances in Neural Information Processing Systems, 35:4475– 4488, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.356799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.356799Z digest=sha256:e8e86b7c68828abf3dff6cab8b0ae03c5131f999bd712fddfe5e86a0d5df2130

Observation 1e388151-715e-4228-b3cf-11399f60e860 · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

BAQ: Efficient Bit Allocation Quantization for Large Language Models BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.534133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.534133Z digest=sha256:c701fe02a866ce3da66e5cf999e7da51a1289aeda1b6ee78d4c3e0c41cdfa53e

Observation 767b367f-c3e5-41fd-9517-9335576483a1 · outbound

This paper cites Gray and David L.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Gray and David L

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.715317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:53.742757Z digest=sha256:c8198c02aec9c41b14e632db29e10a8cdc3517a607200092222981616a8871df

Observation a549fd08-e043-4463-a65f-5907fc04c75e · outbound

This paper cites John Wiley & Sons, 1999.

BAQ: Efficient Bit Allocation Quantization for Large Language Models John Wiley & Sons, 1999

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.878901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.878901Z digest=sha256:92e2bdac1be03137678a5242b755dfdda45c63c9f9c3d9c4ed0b37932eb05894

Observation c8d1f918-2e0c-40cb-aefc-0f4f22b8207d · outbound

This paper cites Jpeg2000: Image compression fundamentals, standards and practice.Journal of Electronic Imaging, 11(2):286–287, 2002.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Jpeg2000: Image compression fundamentals, standards and practice.Journal of Electronic Imaging, 11(2):286–287, 2002

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.352361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:54.023364Z digest=sha256:319ba5bf6f1f992be57c3a9ca0f11f74b69e67a6706492f7d4278506fc9b1835

Observation f486d404-4f12-4790-b780-fea2a955d48d · outbound

This paper cites Springer Science & Business Media, 2012.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Springer Science & Business Media, 2012

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.104520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.104520Z digest=sha256:59390b83443a4fe3ec1b4e52c8bdce0b652a35af952d347337b665a9779598d3

Observation dde2d55b-4767-4ea0-bb7d-46e6631a94c3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.262302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.262302Z digest=sha256:5da3cee70fd089ef2e43ce7072edd69e69209d8a022639a1c5a07ce037d18b64

Observation cc1f0635-fd71-45ba-9a48-a4c83394244d · outbound

This paper cites an unresolved cited work.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.443826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.443826Z digest=sha256:834761886837a49996f89f0aa90a6cf296cd80fed3c695a167c59cfc4620f6f3

Observation 403ab597-f344-4cff-9ee5-8511a7f80c6f · outbound

This paper cites Pointer Sentinel Mixture Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Pointer Sentinel Mixture Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.621632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.621632Z digest=sha256:76654a11e626f2b6d4a3c30c7647871eaa420203c2481b3b5dc0ad7babfedf41

Observation 34702f25-e960-4259-8b8a-e4a69673bc55 · outbound

This paper cites The Penn Treebank: Annotating predicate argument structure.

BAQ: Efficient Bit Allocation Quantization for Large Language Models The Penn Treebank: Annotating predicate argument structure

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.082435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:54.807264Z digest=sha256:376f7b7c231648a463b09f496aefdc6bc58ec7a9de1b1a4710d182c32eee6764

Observation 8e0adfe8-8572-40f5-a731-c30a19440f50 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

BAQ: Efficient Bit Allocation Quantization for Large Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:56.608156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:55.087121Z digest=sha256:a41ab2326a4abd90abdd165d47665e5f144b9f54584d2ec2a1e9e2b521f9992c

Observation 3d4a6eec-b3dc-41ef-84eb-9abcfb0959b8 · outbound

This paper cites Piqa: An algebra for querying protein data sets.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Piqa: An algebra for querying protein data sets

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:56.247548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:55.236048Z digest=sha256:693f372ca6c1c5a68e1f1f71894a4bd529aefd79562477cc1f69f2c7f3310066

Observation e378cbbd-1a98-4b0f-9598-387bd629d825 · outbound

This paper cites A systematic classification of knowl- edge, reasoning, and context within the ARC dataset.

BAQ: Efficient Bit Allocation Quantization for Large Language Models A systematic classification of knowl- edge, reasoning, and context within the ARC dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:55.902652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:55.423637Z digest=sha256:921fb0854fce3635326493230a021435f077508acf03c16b43fc6d458db21605

Pith citing papers

Observation 196ce1c7-1a55-4e9f-98e8-714d0a7fa3c2 · inbound

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs cites this paper.

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:10:24.851205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T17:08:05.945757Z digest=sha256:90fe30ad8ad270d23f7e4c4cffe81ff410da42ecb5636ad65a48dcf63eb59dc8

Observation db559af4-13dd-4bac-9c2a-f077de7e7051 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:40:59.669846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T01:09:21.811344Z digest=sha256:1e4d27ae643d31cadeddfae50ab1121788317579bf37332eb6b80a4238de982e

Observation 6267c1b7-0ebd-46ce-94bf-35493f06b28f · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-05T05:50:43.592486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-05T05:41:13.869451Z digest=sha256:fc8745bd0c550211c57a3243228c7756002ab6bfdf20e1b8b37c911ceea710df