Pith. sign in

Paper Citation Record · LEDGER

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2504.21553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21553 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.928814Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:57.976654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:09:58.199009Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45c07fd1-c0f4-478a-bab6-089220b261ec · outbound

This paper cites Advances in Neural Information Processing Systems36, 34278–34294 (2023).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36, 34278–34294 (2023)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.234203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.838529Z digest=sha256:37883de532b7ec7a3bd2e13b79e00a57cf281f408a2656e1a783cf561e610cde

Observation 8f4fde32-7388-4520-90fe-067769996c82 · outbound

This paper cites Understanding and Overcoming the Challenges of Efficient Transformer Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.843611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.843611Z digest=sha256:b3c57b01e358c68526ff84efbf1c50c9180bce049f3827f6a94a05cdcee9c111

Observation eda4d2bc-0b03-4630-9a0f-72709eda8d96 · outbound

This paper cites int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.848749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.848749Z digest=sha256:67e0d0aeb99636d454290b74a1beaacb68b930d35172002069dc8f1267e957ad

Observation fe0c0ad9-39a8-4277-802b-8f5a05ec22ae · outbound

This paper cites Advances in Neural Information Processing Systems36 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36 (2024)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.853226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.853226Z digest=sha256:3b700accde58fd0a2c372d6fee1393b0ca7aae6f5fd67189f2dd8648dc3615e8

Observation 337645d5-2733-468e-8c36-8c2b374579cc · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.201046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.857680Z digest=sha256:f735c1348afcd8b6f3b092a186fc7f132e18fbca56e79fb47cdea859cb247d2d

Observation 1ae82f16-f726-40cd-be1f-3b2282f3fd38 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.861936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.861936Z digest=sha256:5bc326ede0b5302ef8a72533ddcc839af1322dc4b5d3ba3d454362a4cd9d51b9

Observation 520e269d-5eea-4679-ab18-969282ee0dcf · outbound

This paper cites Understanding and Minimising Outlier Features in Neural Network Training.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Minimising Outlier Features in Neural Network Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.867019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.867019Z digest=sha256:3865c8b5394e4a274a9bf92255f423b70cea0d91d71efd28945add312903ffd2

Observation ddfe4a3b-c127-478a-a440-5e7577d41207 · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.187308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.872026Z digest=sha256:ce53a9ef5d63a13896ebca05c91ce2cb5ab1158cb5c8c59db87f79d1e06cb302

Observation c35f00c6-773c-4eac-a5d3-9b6bfc21fe71 · outbound

This paper cites Mistral 7B.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.876317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.876317Z digest=sha256:c310c944722e1f1fc0c3cfc12588396ecb64b88da2eb79b31c4de673f3a00bf0

Observation 7c4dff9b-d604-4e10-b11d-1e28b491087b · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.173293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.880599Z digest=sha256:c6a271798c9220fc4b86ff75cc1fda91055d561577c488a10ccb347dc9f6b75e

Observation 6f0aa712-e0ad-4e49-a513-41b0ae1e7d55 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.885301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.885301Z digest=sha256:252a3a65591a553cb72509e381e0dd2a644e16a899ef0e00c5ffb81380bf8e64

Observation 33c4ffb8-f0d8-4383-ae7e-e6cd0eb63ca0 · outbound

This paper cites FP8 Formats for Deep Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models FP8 Formats for Deep Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.890019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.890019Z digest=sha256:dfb444d0c269ecd76db1aa210c05b1db9823938425e4f88d258a9e1c29013267

Observation 6b44d73b-7bf7-42d4-9b94-3d55ac77e13c · outbound

This paper cites Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.894637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.894637Z digest=sha256:9f2806202324291d8e01d9e0d8004e45c32598e53db53a895acc81724c548216

Observation 3cb75199-7fba-4c16-a011-927ca194a9da · outbound

This paper cites Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.158237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.898893Z digest=sha256:41f95ac66da868642b4a76624dbd5425eeac3182fe05365281c0b9875c6a6ff8

Observation 12ebaf82-965f-4019-8fe3-47c6875bdf78 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.903254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.903254Z digest=sha256:7aab6ee50d3d19685977ab924e669873f27cee11424ad209fec6fdf023cef5b7

Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · outbound

This paper cites Massive Activations in Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.907873Z digest=sha256:d74dd00a630f3f59fe0911113403a4d31d3aeee1177266610a3f1b9b2a009cce

Observation c33236ec-3ad5-4f97-8a55-7ddcb28b3c77 · outbound

This paper cites https://doi.org/10.5281/zenodo.10256836, https://doi.org/10.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models https://doi.org/10.5281/zenodo.10256836, https://doi.org/10

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.912149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.912149Z digest=sha256:edc184b672237330c6c4ad69175ce9f7bf3daf48bcf39a727fb50db8d9faf13a

Observation 1ee496d5-c5ee-46a8-9f10-5f9e9b15a31f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.916285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.916285Z digest=sha256:328959edb2bb6a251c724dd5e3eda6d866e0c2381eb32282e75df30b43e5fb66

Observation a59e6372-493e-4fba-9906-368a752eeb44 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.920433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.920433Z digest=sha256:a18a7b8de80e21bd22d71de8c44f6f62bd2fe8a5cf0711b8af16c35757f32ba6

Observation 2c5c0b76-c9f0-4067-93de-6e1e8c64449c · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.144216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:05:12.924694Z digest=sha256:8ec9009ae70dac2b4e9f7d847f1d662d50431c48f0ba7f2dd7ef7423e5185ef7

Observation f28bf1b0-3d02-4e9d-a8cb-8af9175b062f · outbound

This paper cites Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.928814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.928814Z digest=sha256:c46400f3c38537229c5dd48a36b50f4c8724e173421e1d703aff403ab8ded3b6

Pith citing papers

Observation 98677ebc-dbe7-41d4-a815-9166739c8b99 · inbound

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs cites this paper.

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:09:58.204110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:09:57.976654Z digest=sha256:c8102f7adec406dd191a4b26cec4d4adfa0ee09f606fcdde36e0f275346b7815