Pith. sign in

Paper Citation Record · LEDGER

A Survey of Quantization Methods for Efficient Neural Network Inference

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 57 inbound Pith citation observations for arXiv:2103.13630.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.13630 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 57 of 57 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:48:57.257985Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

39
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00b66eee-06d4-492e-a75d-7f95a2020df6 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.084102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:0f27eb7d699a02c342669d6ee1d9d8d5801ac885ba2838e3ebb183397a65cb14

Observation 0bba9aaf-d8f3-4125-8ef5-c3afbad10a39 · inbound

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers cites this paper.

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:18:35.220447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:18:35.153078Z digest=sha256:075f2bbefe9e6237320bfec1776a2de023048c640cf356c67e0124f71735bf15

Observation 96bb1c26-5673-40fb-a214-620865aac78c · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.292898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:2600a7b3a7013c024994d420717d383c0cfdcf862538e220bda803bb9da6cc31

Observation b19fa976-c65a-4390-8afd-0e3f0d948751 · inbound

Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity cites this paper.

Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T15:48:57.257985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:48:57.257985Z digest=sha256:13fe51287deae485c4d4ca3ef10f343dbeaa6cdf1dbaae8a25991ad6ebf16976

Observation 27523974-4ea5-4b32-b343-5785644b2cf5 · inbound

Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators cites this paper.

Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:22:57.365228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:22:57.365228Z digest=sha256:2fba2faaf01b56a47847a5495c5df3450d1f78385b7541b793c22ede4f32d297

Observation 409bfe8f-c526-4420-b01d-8342132f3f09 · inbound

Setup Once, Secure Always: A Single-Setup Secure Federated Learning Aggregation Protocol with Forward and Backward Secrecy for Dynamic Users cites this paper.

Setup Once, Secure Always: A Single-Setup Secure Federated Learning Aggregation Protocol with Forward and Backward Secrecy for Dynamic Users A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:06:30.721328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:06:30.721328Z digest=sha256:f46e0a923229af0fd3f3e516f6bd9b1023e275b55070b4ce0cad7f73ca9ecb62

Observation 6de8f12d-cb44-47f3-9775-b6c5d410df9c · inbound

When Do Neural Networks Learn World Models? cites this paper.

When Do Neural Networks Learn World Models? A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T22:12:50.004131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:12:50.004131Z digest=sha256:f8ddc5494601e63b2ccc2d343f35fa3f77c5965131376c072571d4f25dbcc29b

Observation be0c3fb5-8441-4c69-9787-a111ff43e423 · inbound

Can Post-Training Quantization Benefit from an Additional QLoRA Integration? cites this paper.

Can Post-Training Quantization Benefit from an Additional QLoRA Integration? A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T19:03:04.836700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:03:04.836700Z digest=sha256:4f5ce84715a73283fff11e862f940cebd3dd1f0ea963d5600ceada1450f259f6

Observation 99284c22-70a6-4e24-aaa0-5b6928c39c63 · inbound

Forget the Data and Fine-Tuning! Just Fold the Network to Compress cites this paper.

Forget the Data and Fine-Tuning! Just Fold the Network to Compress A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T19:04:45.538462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:04:45.538462Z digest=sha256:1dab72778c8b4a9916d213805cbec8e39f7ed0932aa9815f22e4fffebf40f970

Observation 769f1d7e-efac-45a7-9a6a-5e7e5b8f84a3 · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:01.468760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:01.468760Z digest=sha256:385a7222572b37ca395cc72b5edc7f15916c635b21e707d559844c9bd20c71f2

Observation b2f43a7e-a1c9-427f-89d8-214f9fb73466 · inbound

Large Language Model Meets Constraint Propagation cites this paper.

Large Language Model Meets Constraint Propagation A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:56.274457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:56.274457Z digest=sha256:68fdea7687b2911f4ad6d106373e6ddecbb44fa3507a548f27c5051fca63da1c

Observation e365ab8a-b784-4406-927b-dea0195e3e40 · inbound

BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing cites this paper.

BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:15.911425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:15.911425Z digest=sha256:5368cf8bcaab05382be6d7067fdc0d45dcc6519d929b7cbe0b5545f9406395f1

Observation b9177401-69f6-48b0-bebe-ad54a5599369 · inbound

QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference cites this paper.

QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:39:19.804559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:39:19.804559Z digest=sha256:3534feb31b624b204a86f0e51191dc282bcfcadb56dbf61b8e8832e4ce8ec155

Observation bc07b9b3-bef6-4e53-bf11-8ff2c6c5b35c · inbound

TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents cites this paper.

TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:20.455721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:20.455721Z digest=sha256:c522830e88c734748f61e4d7a67ab0b5e135bb0b3b655c979c48b88475548541

Observation 2b3591ea-676e-410d-b313-49cc164bc3d5 · inbound

Design of an Edge-based Portable EHR System for Anemia Screening in Remote Health Applications cites this paper.

Design of an Edge-based Portable EHR System for Anemia Screening in Remote Health Applications A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:35.654496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:35.654496Z digest=sha256:f14298b63c7127e77117693d1c9b5d77b67265cfdf2b0d31d9e41b1e90e2637d

Observation 30e36c17-7ef8-4375-b83d-e0fc10e3d34a · inbound

Performance Analysis of Post-Training Quantization for CNN-based Conjunctival Pallor Anemia Detection cites this paper.

Performance Analysis of Post-Training Quantization for CNN-based Conjunctival Pallor Anemia Detection A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:53.246566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:53.246566Z digest=sha256:52c2f7449e2ca07d06d9c4e26b865c4cb294bfbf64bfd5589ba0c4c6aa822266

Observation cef6eecd-6b06-45af-8117-5ad273eab734 · inbound

The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm cites this paper.

The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:52:00.477563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T02:47:57.446718Z digest=sha256:90947a45e0b78a60c39d1835385c72107fed49b08e5290a3138a1c2d49a33688

Observation ff657a65-67a1-40a6-87ef-5581be772dfd · inbound

LCS: An AI-based Low-Complexity Scaler for Power-Efficient Super-Resolution of Game Content cites this paper.

LCS: An AI-based Low-Complexity Scaler for Power-Efficient Super-Resolution of Game Content A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:35.904337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:35.904337Z digest=sha256:c689f7bc90b22ae8400ce4728bc3d3e52e6262197b435de050ef2bc92c8c61dc

Observation 7899aa98-c1af-4207-b559-19544490283b · inbound

DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling cites this paper.

DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:11:46.619871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T19:09:04.217591Z digest=sha256:b127a036f921b36f628136c219d04cfea50e2ef30e0ce1982cce5f23a5c8d8a2

Observation 9c2f9e14-7e9e-4bd1-ba72-7da11ad5a3c8 · inbound

Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems cites this paper.

Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:15.380860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:15.380860Z digest=sha256:0ce1de5d44fb88ddaa64cb425e296fe09d021f1b0e20790a2f8c439b8745e65b

Observation 8e164972-68af-458e-a1a3-67e8b39dc80c · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.807509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.807509Z digest=sha256:df6db52ee9b3f5e1486b9c7ab90a38c27fd6b476ad034c451a3dad3396d6ea33

Observation 73dfb776-9b0e-4bf7-ae46-d02998c8e665 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:50:30.316237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:aad1b95430adbbff334abfc0bd0094ae389aa500581b5c2b87b112e8ba764431

Observation 30bb7e37-dce0-4421-a3cc-0708ff5cca08 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:21:38.316714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:21:38.316714Z digest=sha256:45684c60f985d9eaa48f18df76fe5c9520c7c1051564ea952d1a6551af8d3fb3

Observation 25bd78b3-a95f-442e-be59-9f3ed5ccdfc2 · inbound

MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM cites this paper.

MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T21:53:19.873410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:53:19.873410Z digest=sha256:f23eb6606121828bb0ba2fbf2f3483beb470b7eab7622014164a87ba44d0e32e

Observation 1a9ff4f1-1a30-43a5-a7a7-bef9ce394417 · inbound

No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason cites this paper.

No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T19:02:36.023080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T19:02:36.023080Z digest=sha256:362bf1ef419ca6977bccb98dac38ce5fcd5296e9a4cc53aaac43881674c73620

Observation 19e7a59f-1396-4116-b23e-28ada2ad86db · inbound

PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation cites this paper.

PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T22:32:36.129552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:32:36.129552Z digest=sha256:ae7f6126b257493f660755c2169e7ef3cf7863f320ba694f09cb674ee8fb9308

Observation 066bf084-a63e-474a-aea5-f0ea9bbdc345 · inbound

Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence cites this paper.

Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:29:56.944260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T09:27:34.204972Z digest=sha256:965b68b9277d26b50fc385745473c2e03b48c359595a31bcc622995347552fcd

Observation 080d1c6c-d53c-421e-93ca-36bd7330b2e5 · inbound

Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization cites this paper.

Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.645303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:37:31.175385Z digest=sha256:e566249a9ae4b6ef8cf2151d8545ef7f484c7f540c18d1c59dde025a8a49d204

Observation 6e680e32-1d81-4a5c-8cfa-5ff6f0284360 · inbound

Harnessing Photonics for Machine Intelligence cites this paper.

Harnessing Photonics for Machine Intelligence A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:01.129306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:08:41.680567Z digest=sha256:9c620b3fdde989ba37cbacff5e1d19b0fa87eb7a572028d2a30ca0db5b85372a

Observation dc96319b-6373-4a8c-8ccb-f9a11cdc5e3a · inbound

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models cites this paper.

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:08.760148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T10:12:36.972813Z digest=sha256:21d0753b0ced4caaf7f326442fcc88d5839b983dc5976a4915fb767597f17b9a

Observation a9081ec2-8c24-42f3-a64b-926acc3813da · inbound

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models cites this paper.

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:25:46.077628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:12:52.038253Z digest=sha256:66518015f506174e684d5909c4b9655eca48ec1c77618b63a90eccf1a54b2733

Observation 11f0428d-b44b-4eae-aa50-919b9acd2c9a · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:28.021916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:33:41.411292Z digest=sha256:9305116f5880931506bdb1f18b656da3d1eef10e462fb0b0822c486d6eeb8592

Observation 68ae93d3-3789-47fe-a550-34bf64d79dfa · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:59:46.145486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T04:55:01.973832Z digest=sha256:28703143ee57213c81e89f16b168625e6993b5da4ebc54c14a1e5a546b48ffd2

Observation 33becbf7-8bbc-42e0-ae84-ab17dcf3a61e · inbound

Rethink the Role of Neural Decoders in Quantum Error Correction cites this paper.

Rethink the Role of Neural Decoders in Quantum Error Correction A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.516333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T04:54:29.757965Z digest=sha256:a7a965f506493b1480edc2233c2b4651c60ca60d0410cbc04668628eba162d01

Observation 9e9d16aa-7485-4c15-940a-1a1cd0621b44 · inbound

Memristor Technologies for Dynamic Vision Sensors: A Critical Assessment and Research Roadmap cites this paper.

Memristor Technologies for Dynamic Vision Sensors: A Critical Assessment and Research Roadmap A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:47:32.266103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T17:45:42.414771Z digest=sha256:8ca4ad10598ca5a7c80fd71490162d62c49ded60d35da762c1fc1da36c9cfe78

Observation 7c917496-7944-4c49-b644-dd69d6daa2fe · inbound

Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis cites this paper.

Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:29:00.012599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T20:24:16.373261Z digest=sha256:f88655d6c50471025e651f864e5dfdd4d99765cd1dba48013ca1937f05163944

Observation d9ab45ee-97aa-4a14-996e-4c02b83cc323 · inbound

When Bits Break Recourse: Counterfactual-Faithful Quantization cites this paper.

When Bits Break Recourse: Counterfactual-Faithful Quantization A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:43:22.007935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:43:18.182586Z digest=sha256:19628556f58a6dc858834171536d030453d0b5ae79b2a2dff5ddc29a25b323ca

Observation fe01c9bb-2751-4762-b364-575ea68a0c57 · inbound

When Bits Break Recourse: Counterfactual-Faithful Quantization cites this paper.

When Bits Break Recourse: Counterfactual-Faithful Quantization A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:20:53.530912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:20:53.530912Z digest=sha256:af2c2778fd270a175a9e0a88f5338f2f0ef54bfdd30caadb84805c9df4b7fa6f

Observation c0baae1a-f6f6-4ef4-8e4c-bb195ebce903 · inbound

When Bits Break Recourse: Counterfactual-Faithful Quantization cites this paper.

When Bits Break Recourse: Counterfactual-Faithful Quantization A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T05:08:29.889425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:08:29.889425Z digest=sha256:d94fa6b76855a8d9625ce8755c9863b11a6363bb857b1e40f0014fb0bfdbed31

Observation fabdbb09-f437-40e2-b345-ddb832b4111a · inbound

Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation cites this paper.

Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:38:10.880108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T09:34:59.995475Z digest=sha256:c61c8a28fd8af11f39fb19a8ee791c8fd108d6731b0baa5ef60a67858499da45

Observation e53ca334-5dc2-4c5b-9a78-e4134d30fbb3 · inbound

The Thermodynamic Costs of Simple Linear Regression cites this paper.

The Thermodynamic Costs of Simple Linear Regression A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:03:23.273800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T07:03:21.987679Z digest=sha256:5becc744ca5773af5dfd80b008614c9b29c14697863de4abc2b8b95f0fbbaf85

Observation 2ede1385-f556-486c-94a0-9f59a642baba · inbound

K-Quantization and its Impact on Output Performance cites this paper.

K-Quantization and its Impact on Output Performance A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:04.928558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T05:37:46.225362Z digest=sha256:09d857689b4016b978afbd30adb61358ccdc5a5952eeca04316a7b1a00c4fc36

Observation 9e3657d0-edf1-4134-9530-827005f66567 · inbound

Machine learning enables experimental access to photon-by-photon arrival times in scintillation detectors cites this paper.

Machine learning enables experimental access to photon-by-photon arrival times in scintillation detectors A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:53:17.409156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T09:49:07.517164Z digest=sha256:7a41c7fed2431fec10f496c879416028a04c971478f318bb978f0d78b7ca834b

Observation 2ef426c2-1881-4425-84a0-1cb010ea3651 · inbound

The Complexity of Verifying Feedforward Neural Networks in Quantised Settings cites this paper.

The Complexity of Verifying Feedforward Neural Networks in Quantised Settings A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:49.989139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T00:09:07.981665Z digest=sha256:444528be3f0c20fa54cb29cb5acc769d0128723e73002ce950772cca9ff383f7

Observation 8d05e762-cfbb-4f92-a600-89c920cfaefb · inbound

PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference cites this paper.

PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.534010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T15:08:25.880424Z digest=sha256:30135f7edc7ab7f712f76a625155859f65492ac3876ab31b26fc0986228e9ed4

Observation df96f0e2-f023-4824-9b30-b2d5de7db15d · inbound

Minimum Distortion Quantization with Specified Output Distribution cites this paper.

Minimum Distortion Quantization with Specified Output Distribution A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:37:45.441499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T11:58:15.793011Z digest=sha256:44947fb9218900b51395d268ce9243e907419aed29b70fc26b60fda863ee561b

Observation d6802a1d-cf49-44ce-bbd9-27ee1d7de2a3 · inbound

Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score cites this paper.

Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:58:22.016841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T07:21:21.103554Z digest=sha256:2e863c1cb6c9e0a15abff760c506875faedd274c7685aff4cf7ee4b98e0b612b

Observation 77947d34-ba98-4087-93dd-ceac11290e4e · inbound

On the Expressive Power of Weight Quantization in Large Language Models cites this paper.

On the Expressive Power of Weight Quantization in Large Language Models A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.272097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:53:45.787243Z digest=sha256:c076dae8782db0acdc238dd22ea7bf354ac4235307000646030ba3f0e2c0134b

Observation e5514e51-a7fd-494b-9431-cf45e6db27f9 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:59:52.403341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:1633455a9ae9f0ebb027b92b98215c4b55024d49ab290abe01bf540c4a0c8fcf

Observation cfe547c6-22bf-441d-a7c8-09ba03676a89 · inbound

Quantization in Federated Learning: Methods, Challenges and Future Directions cites this paper.

Quantization in Federated Learning: Methods, Challenges and Future Directions A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:51.289218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T04:56:12.424439Z digest=sha256:42a957cddec8cd5ad0b8764284ecd14d18ab972f49c60c24b31aaf2d79d7dfbf

Observation f5d6647f-9421-400f-aee8-8bac83eb7ee4 · inbound

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space cites this paper.

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:55.881300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T12:35:58.613973Z digest=sha256:d4fa70e1dd05fc186667fe9e1d861650094d1bb92fc3f3c62d53d0b461af0441

Observation 4159fcf2-36dd-47b0-80b2-a690665a03d8 · inbound

Lyapunov-Guided Training for Hardware-Safe Neural Networks Under Fixed-Point Arithmetic cites this paper.

Lyapunov-Guided Training for Hardware-Safe Neural Networks Under Fixed-Point Arithmetic A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T17:52:29.131365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:52:29.131365Z digest=sha256:7dd97f7b9ea09ce8f9e0362ad1b83ee2f196a9335f2a84b351be87a1a98e09f9

Observation c0f57189-cecb-468d-a4fc-06c51b67d633 · inbound

ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM cites this paper.

ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:37:47.281341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:37:47.281341Z digest=sha256:7307cef7fef7fd1047e475f6660874c5badcbbdc940af70058c0739070c7a566

Observation a6a1969b-5349-42d7-9add-41d6457898bf · inbound

QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs cites this paper.

QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:27.328655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:27.328655Z digest=sha256:9da90af7cf99e61d7aeab5a17d40d41564e19496060d96eab5ffbf4851506ec7

Observation 9d56633a-8464-4bf9-b762-95bb344dd067 · inbound

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment cites this paper.

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:14.757587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:14.757587Z digest=sha256:4135690975defa2ec76f509a18e4ef287bf1fe074046edfabd95a6ed238593b2

Observation 824d8568-6119-4bcc-91c5-30e341240d80 · inbound

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers cites this paper.

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T02:57:27.356357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:57:27.356357Z digest=sha256:a4b5728b1963500774ed10d335b718d001c0134a6f083668b0945df39f8fe43b

Observation 04db17ea-4882-4970-9c26-d5fb26d0e2c5 · inbound

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models cites this paper.

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:45.725654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T01:04:45.725654Z digest=sha256:895cdf48db853d48f7893317840d9df8d00a7489068b9f283fd41bf0f164b018