Pith. sign in

Paper Citation Record · LEDGER

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2504.21553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21553 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.928814Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:57.976654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:09:58.199009Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45c07fd1-c0f4-478a-bab6-089220b261ec · outbound

This paper cites Advances in Neural Information Processing Systems36, 34278–34294 (2023).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36, 34278–34294 (2023)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.234203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.838529Z digest=sha256:2f9837bbb02dbe3ec50ef315c4f43bb5b29d4208063c2b97b42215cedf66ee5e

Observation 8f4fde32-7388-4520-90fe-067769996c82 · outbound

This paper cites Understanding and Overcoming the Challenges of Efficient Transformer Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.843611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.843611Z digest=sha256:7ceb0828a5969d22e724cc2f0bdace1610500f666a0594c31cd864f48937b970

Observation eda4d2bc-0b03-4630-9a0f-72709eda8d96 · outbound

This paper cites int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.848749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.848749Z digest=sha256:a8dd80958d49e21a592586b7a30a65e97ee2d2cd77277bd30010b189efaa1feb

Observation fe0c0ad9-39a8-4277-802b-8f5a05ec22ae · outbound

This paper cites Advances in Neural Information Processing Systems36 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36 (2024)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.853226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.853226Z digest=sha256:902e560f5f280e26418701facd1a7afbc44969c99a51aac3ec76315a1da86a1f

Observation 337645d5-2733-468e-8c36-8c2b374579cc · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.201046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.857680Z digest=sha256:1f60757439455e08c20291215d2fe32b39e14160fbdd6649c1a0c2501cf02c80

Observation 1ae82f16-f726-40cd-be1f-3b2282f3fd38 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.861936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.861936Z digest=sha256:d53e3a9156e33e463ea0a0799af588795f185febe3433c7eb5beb658e6e1e09b

Observation 520e269d-5eea-4679-ab18-969282ee0dcf · outbound

This paper cites Understanding and Minimising Outlier Features in Neural Network Training.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Minimising Outlier Features in Neural Network Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.867019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.867019Z digest=sha256:164768ad98a519e69f0b55c61328f03dd6d0d9a2509fa44847e8d35dda61d531

Observation ddfe4a3b-c127-478a-a440-5e7577d41207 · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.187308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.872026Z digest=sha256:ff3dde9e206877e29221dfaabb195eb1306e26c5661bea95539fba8a49ff8488

Observation c35f00c6-773c-4eac-a5d3-9b6bfc21fe71 · outbound

This paper cites Mistral 7B.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.876317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.876317Z digest=sha256:0554f687beb7c56abcffbb148da2763e94d6b868a27e2f76a59f8c4f50cc9fbe

Observation 7c4dff9b-d604-4e10-b11d-1e28b491087b · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.173293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.880599Z digest=sha256:987e8103ec875fa8a192f5c249f05313cb923aded8ff987399f108e669b92916

Observation 6f0aa712-e0ad-4e49-a513-41b0ae1e7d55 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.885301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.885301Z digest=sha256:53f9b4a0b7c0b63f05d194435153ae3faa2bea08e1162942dec083827aae8809

Observation 33c4ffb8-f0d8-4383-ae7e-e6cd0eb63ca0 · outbound

This paper cites FP8 Formats for Deep Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models FP8 Formats for Deep Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.890019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.890019Z digest=sha256:3ec8ea7b32a60ce0249a119d5daa8f59317a7f13832af3ca9a640dac833e9984

Observation 6b44d73b-7bf7-42d4-9b94-3d55ac77e13c · outbound

This paper cites Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.894637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.894637Z digest=sha256:17b53329f2749161e231ac798827825db371c145d52067fd9267163c6ff0fb10

Observation 3cb75199-7fba-4c16-a011-927ca194a9da · outbound

This paper cites Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.158237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.898893Z digest=sha256:8f4ba4802ddc0fd5dbd9b2316f8c4dec4ffb37c28154b0a4d6e97e811bca8a0b

Observation 12ebaf82-965f-4019-8fe3-47c6875bdf78 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.903254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.903254Z digest=sha256:787d329e0b9e3bbe012bf077e5351582dc1d17b7e3fed356eae3b43b0c74ee68

Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · outbound

This paper cites Massive Activations in Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.907873Z digest=sha256:21006a21e696f24c4b3b05ab815b51446936a34e4a261b61f4906e1c7fa9acb9

Observation c33236ec-3ad5-4f97-8a55-7ddcb28b3c77 · outbound

This paper cites https://doi.org/10.5281/zenodo.10256836, https://doi.org/10.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models https://doi.org/10.5281/zenodo.10256836, https://doi.org/10

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.912149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.912149Z digest=sha256:b1fe7b55b7e153e7792514b39e644c5b3b54369804c0374147a461fb9a50e557

Observation 1ee496d5-c5ee-46a8-9f10-5f9e9b15a31f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.916285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.916285Z digest=sha256:0d5cea5ac134fca8f36950536c675ccbe3b69aed287ee16be8894867e18f6699

Observation a59e6372-493e-4fba-9906-368a752eeb44 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.920433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.920433Z digest=sha256:3b867e35471a8045ea68ddd43e7f526142d26ec24ce1418b4b6d26a8c6da03f3

Observation 2c5c0b76-c9f0-4067-93de-6e1e8c64449c · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.144216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.924694Z digest=sha256:3496ae95e1afbe6666bb3fa072de4e5c911b08c681b58a5fe725bd8901840881

Observation f28bf1b0-3d02-4e9d-a8cb-8af9175b062f · outbound

This paper cites Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.928814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.928814Z digest=sha256:2389a0aa5ddd127cbd69523e19f2e5be920f8c6c7ef9b5a332d8e15b2bd98ecd

Pith citing papers

Observation 98677ebc-dbe7-41d4-a815-9166739c8b99 · inbound

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs cites this paper.

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:09:58.204110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:09:57.976654Z digest=sha256:2007871b3dcc2b910470980b7f764dece9144827ece7c0ca3ccbd92f3034887d