Pith. sign in

Paper Citation Record · LEDGER

FPTQuant: Function-Preserving Transforms for LLM Quantization

As of 17 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 6 inbound Pith citation observations for arXiv:2506.04985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04985 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:50.212638Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:54:26.612471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:11:14.387607Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved39
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ba8cc1a-738b-4380-916b-0bd6202341e8 · outbound

This paper cites Understanding and overcoming the challenges of efficient transformer quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Understanding and overcoming the challenges of efficient transformer quantization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.626006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:43.392506Z digest=sha256:efe40a8e2e1aeccae26c9ad0f5db85b4ae1dcda6162ae8217115db1e6f2d5a4f

Observation d8256077-1503-4508-a07d-c06095ff042e · outbound

This paper cites Bert busters: Outlier dimensions that disrupt transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization Bert busters: Outlier dimensions that disrupt transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.355244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:43.456213Z digest=sha256:38717223378dd0a35a07b445f9bebfc400ed921577867085343e5ea07fb676a8

Observation 58f5443b-3af2-4d9a-a195-b157c627bcf6 · outbound

This paper cites an unresolved cited work.

FPTQuant: Function-Preserving Transforms for LLM Quantization Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.539598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.539598Z digest=sha256:7158d275555f98c917edfc9e6c8d495252d68d6874f98c0c4be8933960ec9a6b

Observation c7a78b1e-bcb8-4804-99bd-ffe86490b644 · outbound

This paper cites Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.611046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.611046Z digest=sha256:d1796f1bb1ca54ab1f3d6bf591c876d88bae1e09861975ae08058dae7e6fa3e2

Observation dbc63c28-99f4-47d7-b4b6-d19a1ba1c380 · outbound

This paper cites Massive Activations in Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Massive Activations in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.694378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.694378Z digest=sha256:8fd0e6a626525504536074f8f24a23ed672dd8654a53f7c813d256875b50c38e

Observation b8559f51-d5d2-40ec-adc0-7fce677ae39b · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.819233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.819233Z digest=sha256:c9d3901bd733afbf58c8113a28fd40d8aa69c4fb5c7b5c6a29b12eb88aa630f1

Observation 4d5c4acd-5ea6-4cb4-94c9-3db336267736 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.915764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.915764Z digest=sha256:36b68f06a187c7910890f48d29e3a979564ee25286d8b51e86f8a8568d0e98f6

Observation dd047719-78d1-4eb2-be08-62d303645640 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpinQuant: LLM quantization with learned rotations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.996956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.996956Z digest=sha256:d2f8bd8c07a5debbb6f46fbd5a93f4b98abfa70451139a6aa306fbd648cf7f55

Observation 509b68ee-495b-49c2-b78c-d8b23520fd09 · outbound

This paper cites Quantizing deep convolutional networks for efficient inference: A whitepaper.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizing deep convolutional networks for efficient inference: A whitepaper

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.073059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.073059Z digest=sha256:e5a5f0c2a91741b2f3a50941148ca634fd09e809312070fab73ef3c6f329486e

Observation cd97d06f-d500-4c0b-bfb6-2e87bb99de26 · outbound

This paper cites A White Paper on Neural Network Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization A White Paper on Neural Network Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.171780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.171780Z digest=sha256:19a23a75fd7d0f33f23d87e1c1e001a97dd7ffbadbeb44c5851f0d98bf3fc1e1

Observation 3c83e253-9f44-4951-9bd2-dd5921bec1df · outbound

This paper cites Post-training 4-bit quantization of convolution networks for rapid-deployment.

FPTQuant: Function-Preserving Transforms for LLM Quantization Post-training 4-bit quantization of convolution networks for rapid-deployment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.249922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.249922Z digest=sha256:b4a6370d2952d6595855883bfdd198687f407ae892a40c4cd22bbf9b408f8f85

Observation 50e04b74-1f22-4005-b250-0dacbfd07d84 · outbound

This paper cites Zeroq: A novel zero shot quantization framework.

FPTQuant: Function-Preserving Transforms for LLM Quantization Zeroq: A novel zero shot quantization framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.104781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:44.317595Z digest=sha256:2fe6eaee613cce3b9e11e9e94f48a9771c9ec4380cf8df78aa828005a4c630bd

Observation 885d442f-996e-402c-bedb-82def12ce75b · outbound

This paper cites Low-bit quantization of neural networks for efficient inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-bit quantization of neural networks for efficient inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.881070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:44.407583Z digest=sha256:0f9a38751addfcad97f203d06f75966f616d9f60d71c7e97a023812b0bc94a91

Observation 438057e3-01be-4626-a686-f32d716e6d88 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.490409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.490409Z digest=sha256:aa112f9e5a68c4f7327da2a0b6d9224321625213994d0f591e22d02a0df4a3b3

Observation 1f216aed-f293-45eb-a214-bb2a4a85655a · outbound

This paper cites Same, same but different: Recovering neural network quantization error through weight factorization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Same, same but different: Recovering neural network quantization error through weight factorization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:44.585531Z digest=sha256:553785ed12cd130322165122921149eb0ced20e130a8a9761aa06e27f6495b67

Observation 6b191482-579b-4f32-8e53-101cfa8e6b23 · outbound

This paper cites Improving neural net- work quantization without retraining using outlier channel splitting.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving neural net- work quantization without retraining using outlier channel splitting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.392159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:44.660979Z digest=sha256:9e21696de6676ed1e40e470a9800fae0c8a720909ae9d0161591e7ddd44abe8c

Observation 2aed43fc-d301-4c5a-917d-c633adb95ae8 · outbound

This paper cites Data-free quantization through weight equalization and bias correction.

FPTQuant: Function-Preserving Transforms for LLM Quantization Data-free quantization through weight equalization and bias correction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.110221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:44.761658Z digest=sha256:def565f17014839518be6427fe4a5dffdb7c82a768156f7016ff99497443988b

Observation b210edd9-11bb-41f6-af73-0a1ab55e1632 · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.852414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.852414Z digest=sha256:6211bf72759c32d719fe25cceeb47194f5f4eb3685a53ceb590c6da74dc88618

Observation e388f3a5-d0be-46bb-93b9-73a199e95015 · outbound

This paper cites BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction.

FPTQuant: Function-Preserving Transforms for LLM Quantization BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.932932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.932932Z digest=sha256:060bf9ee88c8ba71e56470b4fc490669664cd33072ae365a5287d9d51cd15964

Observation c024df58-8b6e-4c68-be7e-3940bd5d7b7c · outbound

This paper cites Deep learning with limited numerical precision.

FPTQuant: Function-Preserving Transforms for LLM Quantization Deep learning with limited numerical precision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.882906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:45.012514Z digest=sha256:cead3fa92343f3b543705a364258f401a7e9555a44ce31a2af4364bb2c517644

Observation 5cc967d0-211a-47cf-9b83-fad6bcc52476 · outbound

This paper cites Quantization and training of neural networks for efficient integer-arithmetic-only inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization and training of neural networks for efficient integer-arithmetic-only inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.735484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:45.059061Z digest=sha256:e332822284e0e54ad3e90604cbe55277ecbe25b44065232ad13e6452822d18a8

Observation 969ea878-426c-48d8-bc60-02170c14f88e · outbound

This paper cites Esser, Jeffrey L.

FPTQuant: Function-Preserving Transforms for LLM Quantization Esser, Jeffrey L

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.535555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:45.133526Z digest=sha256:c733fd429806ef25fab0f98edeb3cab5a49c563b3ed31747315b09e0ce44f2b6

Observation 485f1515-d3bf-4677-986b-9f33a9ddafdc · outbound

This paper cites Lsq+: Improving low-bit quantization through learnable offsets and better initialization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Lsq+: Improving low-bit quantization through learnable offsets and better initialization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.246016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:45.211084Z digest=sha256:5b2c9c1af83ea302598f7bfd5a5b2689b0c31c5018899d128013a0c16ee3001e

Observation d507e2b7-65cc-4069-b782-597c119b6fd5 · outbound

This paper cites Overcoming oscillations in quantization-aware training.

FPTQuant: Function-Preserving Transforms for LLM Quantization Overcoming oscillations in quantization-aware training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.037723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:45.287584Z digest=sha256:2837ce3be50d1b47e27c9c76de00c40d249a142da2b1f5af7b5f42641917a314

Observation 7da8cef2-4587-4574-81db-7b5309ac1165 · outbound

This paper cites LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.412441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.412441Z digest=sha256:5ba3a64417a96ccb0ce349ac4efad2066f250e80eea61b54985a96fc7721fd9d

Observation 3cb6aee2-6a07-482b-b315-3ac18a450963 · outbound

This paper cites BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation.

FPTQuant: Function-Preserving Transforms for LLM Quantization BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.503765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.503765Z digest=sha256:4e50e749ba0622a919d4d3f490e4927e45f8a4260c5bf5c7f80bb15f86b04d36

Observation f030abb1-84cb-4181-94c3-7817bbfdee0e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.644520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.644520Z digest=sha256:847334ea5e11f5acaaf5faef4b91bb6252e0d2646651947419d7817d54cef798

Observation 657cfb93-0532-4c48-87ab-aec4df2bec6d · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization Qlora: Efficient finetuning of quantized llms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.723371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.723371Z digest=sha256:9ab1de7bd687bd429b1148542914602c3613360ba5530db07fe22ee0c171e60f

Observation 289ff14b-bc35-4733-ae33-54f1fb574b65 · outbound

This paper cites QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.821751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.821751Z digest=sha256:b969ee2997a226eb979e8b86df0d9d08391c614256b70d67ca2a3586d2c6a5dc

Observation 1768de1f-a55d-48c9-be05-903aee41c92e · outbound

This paper cites Low-Rank Quantization-Aware Training for LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-Rank Quantization-Aware Training for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.907542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.907542Z digest=sha256:5a9c6b81e3dac088b0f090a26487b056ebd2fa468f50aefcc98130b78b260a47

Observation e4a6e479-b504-4dfe-b7b0-2a07dfdb3d53 · outbound

This paper cites Paretoq: Scaling laws in extremely low-bit llm quantization, 2025.

FPTQuant: Function-Preserving Transforms for LLM Quantization Paretoq: Scaling laws in extremely low-bit llm quantization, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.985375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.985375Z digest=sha256:01e6e7b0ac29a02a8161956e285e3fd29c0ae3c594a16d5bcee27881d7401efa

Observation 15fd83d1-ec0c-4244-85a5-e58b9eb4ac3b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.065076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.065076Z digest=sha256:1ce8928d259494206f8fc653f829785215d8a20679a48c516c8c5878d4b2c882

Observation 9ecc4d88-0f0c-4804-8227-bdda6b9b0470 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.139211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.139211Z digest=sha256:c4af5c38b6e4847f0376d7a72d305e8ecc7d387aed7e062e7648cad43732da6d

Observation 8fad7b8c-9a67-4e89-8b00-bbee994ae2b9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

FPTQuant: Function-Preserving Transforms for LLM Quantization AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.226301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.226301Z digest=sha256:3ae7e16324c318a99fb9d04f90786b31dedd773529d438f1a68581d9f9c733de

Observation c502772d-b90a-412c-8437-8958ce0fba68 · outbound

This paper cites Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.309944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.309944Z digest=sha256:846eda1b59a3d082d2027bfe28d89a2f7ce8d63cfddad54239934e2fad3b28db

Observation 8667e534-2518-4bf2-954d-47bdcf62d9c7 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.395454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.395454Z digest=sha256:1b1f83707f2fe5ac6436ef02733bce52cc59a31c2df2ad35579d708fa6665430

Observation f696d429-04b3-4cc4-a2a8-feec06fcb35b · outbound

This paper cites SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.491664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.491664Z digest=sha256:bd26422deae67fd24099e831a116d77c81e39fa4f8b21f09792af603068fdc44

Observation 9f794de9-c001-4ba4-8925-f3957faa6e3a · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Extreme Compression of Large Language Models via Additive Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.566561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.566561Z digest=sha256:e482c1c22e713cba1d538f55104f3459bd6020a3feef7467d968f72e55da0eb7

Observation ec96620d-493a-4815-8591-81342b0588be · outbound

This paper cites A frustratingly easy post-training quantization scheme for llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization A frustratingly easy post-training quantization scheme for llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.828721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:46.675630Z digest=sha256:2eebeb5b36332e9889bf66ccf9497f4054aefdef0e66cffd2d1675676ecb2b4c

Observation da98e275-fbe4-41a6-ad7a-77f7f3f47aa0 · outbound

This paper cites Flexround: Learnable rounding based on element-wise division for post-training quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Flexround: Learnable rounding based on element-wise division for post-training quantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.667231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:46.752521Z digest=sha256:2b916fd460f1c772f8028494fdc3ba1c9561ff41cbc991ae2b4be8a8944e39e0

Observation bf7d682b-5699-4ec9-bb14-7d3682149f37 · outbound

This paper cites Long- range zero-shot generative deep network quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Long- range zero-shot generative deep network quantization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.454601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:46.841044Z digest=sha256:4a83bcfb8f6fcaf51d661779e306a77f5df6422b61074dd91e3f8ab5e3ca6044

Observation 0a53fc79-93eb-4c1d-989b-0d38e74703a7 · outbound

This paper cites Quip: 2-bit quantiza- tion of large language models with guarantees.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quip: 2-bit quantiza- tion of large language models with guarantees

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.223241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:46.918388Z digest=sha256:ca67cd327d3ac70eaafa62ed779b96912570624f974d0c0bcfa36a8fef8a8a36

Observation a2cac9bb-05ea-4ff6-a91f-d0c1ecd524df · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

FPTQuant: Function-Preserving Transforms for LLM Quantization Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.980855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.980855Z digest=sha256:a4fe747e2e55053cbc52a765b2328d1e4099f2d5695c6d3427ef559c8884d8e6

Observation 69b59f67-87e3-4053-98ab-73f54a6393f9 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.032509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.032509Z digest=sha256:ea2fe0ccd7747381142816e9fdd9c6c6887ac3fe12de012d978bdaf821eb9daf

Observation e26850cc-f1fa-42f3-8bd4-b7d17fd82973 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.975213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:47.124919Z digest=sha256:29a46a6ed4e17f6cfb60dc01a0e3220b811249c72152ce6a13f78563a3bab4c4

Observation 870a263e-334d-4d85-bcaa-7fa4cc1a2161 · outbound

This paper cites DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.305676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.305676Z digest=sha256:3e5a9796c2f5140ba0bc21189431c9c267cce74f01bb418468fe6a866c381339

Observation 252b32eb-f7f4-4bcd-809a-0b2ae9fc5dcc · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization FlatQuant: Flatness Matters for LLM Quantization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.377076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.377076Z digest=sha256:ee501891f77146d873861d92afce3c4a452a1fe3df14c96faf0b52ad8067e99d

Observation 64dc9d2b-c2fd-430f-a0f7-a6a009f3ce69 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.464816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.464816Z digest=sha256:faf3400548d756d41cba4df8ba4c8c067d4a39772f7df15e23fbab3c62bf30f6

Observation 7287845d-6e7e-4085-9c68-9ffa36c6253a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

FPTQuant: Function-Preserving Transforms for LLM Quantization Roformer: Enhanced transformer with rotary position embedding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.553216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.553216Z digest=sha256:045f0293d5341364d10fbb4e4779b5dd3702bd7d599b1cbb9e0be5a19862adfc

Observation 43470d1e-cbcb-47c7-af8c-9615542eca02 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.634359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.634359Z digest=sha256:5ec4501f0f6845b6e88d745843e7e5b37b50a206ec9d44a03f78cdb3abf1de66

Observation f00ec514-0be5-4c54-a336-416364dbeb24 · outbound

This paper cites The Llama 3 Herd of Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.714871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.714871Z digest=sha256:a69f18523557459748bb5f987c90654a6c50dc5e9093cdc82c5f3ea52d4865e1

Observation 16655abc-63aa-47d1-bdef-c8b6b42e0071 · outbound

This paper cites Pointer sentinel mixture models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Pointer sentinel mixture models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.806276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.806276Z digest=sha256:babb919ec397a076e88a348c457672224aaef2b50d1aa3caa4bb1b1578fb7020

Observation 19c78207-9ae1-4dc9-996b-c6d7aeeb92ed · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

FPTQuant: Function-Preserving Transforms for LLM Quantization PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.813234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:47.981754Z digest=sha256:f89a99140d82f8b88ef8a7dab6ddd905f862c4b4d095d2312a6177d5ca52e7cc

Observation a34ff82a-8ca1-475f-b07b-e0983c96af3a · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.

FPTQuant: Function-Preserving Transforms for LLM Quantization WinoGrande: an adversarial winograd schema challenge at scale

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.636507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:48.097905Z digest=sha256:03504c681d681043153da6aa20d6432fe6bb105ce85dc848a4fb3bb76345999d

Observation 1d04d0d8-2b3b-488a-99d5-7ca5f7ae42f5 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

FPTQuant: Function-Preserving Transforms for LLM Quantization HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.336766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.336766Z digest=sha256:7e1d54431847400960573fe44ceff591d475c7f2fb6d733e92dde5306961aaaa

Observation e7ad51cb-3060-462f-87fc-3d24e2b28753 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FPTQuant: Function-Preserving Transforms for LLM Quantization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.475734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.475734Z digest=sha256:dd591774144f611665abe2c2074df00a7373df9507083403f4b29088e921c672

Observation f46dc064-b0cf-4035-a0c0-11d2169c7aa3 · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

FPTQuant: Function-Preserving Transforms for LLM Quantization The lambada dataset: Word prediction requiring a broad discourse context

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.430817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:48.616558Z digest=sha256:5fc06d253f23dab3c16045e700766f1abafd80dc5d3221068c48f95ede79db76

Observation f3c35bc1-b422-4358-baea-1ae5efa5ca4d · outbound

This paper cites Working with Quantized Types — NVIDIA TensorRT Doc- umentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Working with Quantized Types — NVIDIA TensorRT Doc- umentation

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.245750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:48.792727Z digest=sha256:bb4e68f4e4c9cce0d88a5c88953cd2a37ecfcc5196a680b3e883f1d51a616a72

Observation e737f547-f15a-4453-a5aa-4dad17c7128f · outbound

This paper cites Quantization — PyTorch AO documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization — PyTorch AO documentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.033620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:48.899498Z digest=sha256:1520e7733683078b351371e71894295e3958b0e173671e20069f3b4544f7d09a

Observation 30abc064-9048-4414-a3c4-459a113e0e72 · outbound

This paper cites AI Engine Direct SDK documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization AI Engine Direct SDK documentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.915494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.036004Z digest=sha256:ef5f69eb1639619e93cbe44b88c0e3b9c095ab659112a6b4118b7041674ba5df

Observation 9d768a9a-4f8f-4431-ba16-962eb3325d66 · outbound

This paper cites TensorRT operators documentation: DynamicQuantize not supported on DLA.

FPTQuant: Function-Preserving Transforms for LLM Quantization TensorRT operators documentation: DynamicQuantize not supported on DLA

Reference 61

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.766879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.215303Z digest=sha256:f4e367b396164600f5a2e5ecffa3ed6d66fa959e8f97939805a24a67575e0ef8

Observation e72a981b-b49f-4efb-80c4-7bc224be7ec9 · outbound

This paper cites fast-hadamard-transform.

FPTQuant: Function-Preserving Transforms for LLM Quantization fast-hadamard-transform

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.513457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.347868Z digest=sha256:41c6d0d18a1b9a2d713a3a2b5c43eb1d344cbd0be519014ab4044b66f238f639

Observation a265f935-0e7d-45f6-9090-3ef70eb62590 · outbound

This paper cites double-packed.

FPTQuant: Function-Preserving Transforms for LLM Quantization double-packed

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.383814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.457392Z digest=sha256:77a6f5ac6a9f76df714a6ba1d64ad7772d234fd63e4a3f082ae093c503bd2e1e

Observation 389e9d3d-8215-4636-88c7-3c78d8a8139a · outbound

This paper cites Evaluate quantization error per quantizer placement (e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Evaluate quantization error per quantizer placement (e.g

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.260528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.606480Z digest=sha256:dea2edd942d19292562f7444e50f78c03fd040ca2aa2178e39c4466c33828a34

Observation 0d901eaf-5cd8-4c22-9a12-499f44daeb7a · outbound

This paper cites Based on step 1, choose which FPTs to add: (a) Attention and FFN input.

FPTQuant: Function-Preserving Transforms for LLM Quantization Based on step 1, choose which FPTs to add: (a) Attention and FFN input

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.990152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.704829Z digest=sha256:e90ee492030faf3574f1e7edf7432c3fb5ab1695bbced0b6b7703b801f948a98

Observation b57a2e1a-8f17-4c26-90bc-d603c2affe6a · outbound

This paper cites Initialize transforms, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Initialize transforms, e.g

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.744777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.859176Z digest=sha256:b72ea38d77a4836d0e6fca0324fb57f8ca942d0efcd241deda00b129a5460675

Observation fa9295aa-18e8-4de6-a6a1-cb6f538260c3 · outbound

This paper cites Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1).

FPTQuant: Function-Preserving Transforms for LLM Quantization Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.540897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:49.988735Z digest=sha256:d1e6404203dfe3e462af879920aebeeb05cf6443103ef89ee7b4382287be5c1d

Observation 33ec7427-9afd-4f78-a2fc-1c9b3d06c178 · outbound

This paper cites Set the initial quantization grid, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Set the initial quantization grid, e.g

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.372855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:50.099603Z digest=sha256:af17b5112c6656cdfad6ac2204a766e7d7485359e4ccc99b3c1c60e79df5fe59

Observation b4549857-6b67-4f37-9003-499b98980c9d · outbound

This paper cites Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target.

FPTQuant: Function-Preserving Transforms for LLM Quantization Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.087441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:43:50.212638Z digest=sha256:f1f9d271db57cf304bddedd191ac393b6857e920c4ffbb4cbbec78c476d5845e

Observation 9a0cd5c9-7041-469c-94cf-c7e320333fb7 · outbound

This paper cites doi: 10.1145/3474381.

FPTQuant: Function-Preserving Transforms for LLM Quantization doi: 10.1145/3474381

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.191111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.191111Z digest=sha256:f40ca81f2aa01429fbdad0b31c49573158b445e1333bf18a06787394137dad4a

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:8eecbf039c832de99947e50062b8e1eb1376bb84349420d630094e227d74fc3d

Pith citing papers

Observation fbe7599b-8d5c-4a5e-b8b4-f20481b37886 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:896cafad07ba505c14e05ecaf4ad22bf8111664f590215794928d8a1f41130f7

Observation 7bd8316e-ec7e-42e8-9b9d-7b05baa4b7ba · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:0cc8d808e12301a4655d0857fec908cc6e60407358b2ee9b6d5ae643f75d27f3

Observation 781beda3-9b01-4397-bf23-090a6a93836f · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:36.313617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:675c2676840f0fb03752e1a64dd16ab64a9433e395c098e8b756628015724f1e

Observation 70b001ba-0d1e-4e7a-8c1e-dd1c65847f4e · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:35:12.908124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:35:12.908124Z digest=sha256:25b1188ec07208edbe85c5c78feff0af12bd253831ac466bcfaf162d3fcb062b

Observation f81c0d0d-901b-418c-bf2b-bff90bf5bbf0 · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:53:34.398685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:53:34.398685Z digest=sha256:bf0b76d5e340b12fe86e8c22d2be2c51875807d5d1b405878549d29459cb2d7d

Observation e06b035f-5455-42b9-a447-a8d42814148c · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.612471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.612471Z digest=sha256:df2d31b941ddea1227ac4af0dba4f89f8c4f89f5c87c6e667e6ce2be6f2a28b5