Pith. sign in

Paper Citation Record · LEDGER

Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2310.19102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19102 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:44:48.486215Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

23
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fbdf8ea6-4f0a-4542-8838-cace19618111 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.259219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:cb008fc6445427cbec6a3c983bea0bd11d8fa6c235ccbcd2fe7d9ba8508b068d

Observation 14b49881-a09e-490f-b4a6-f518a0782182 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.468374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:00e38cbceec971dbbfbb607147c8a3baf100df0d7de6a34c309da000073a3809

Observation 084a6871-9427-4e92-a23b-92f0da99b653 · inbound

SpinQuant: LLM quantization with learned rotations cites this paper.

SpinQuant: LLM quantization with learned rotations Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:52:34.799916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:52:34.606853Z digest=sha256:208e3524f74bd6cb0d19dc1b8053f60869309d00c58259b259ae27de755df123

Observation 7e3f4934-d3b9-49bb-a126-fac795bbde37 · inbound

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations cites this paper.

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:44:48.486215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:44:48.486215Z digest=sha256:70162d07191f02e4621a34276e7410ccbc83f592092ab82a970c4723895d1ae0

Observation da10163d-d38c-4102-ba5e-e01fcc73c35f · inbound

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs cites this paper.

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:47.771772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:47.771772Z digest=sha256:a4c15a4c3c641de0eeb152df76406521caf8869b12f818327956726388183df3

Observation c3432691-8733-4ffd-aab5-4797d67bc395 · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.829325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.829325Z digest=sha256:90ac65b1d5321dbb2d411899db494a4dd10454fddad56d59f4ba7112b0eceb42

Observation 69192696-a043-4bd7-8fd5-10dd46887f6e · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.208524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:f537f9d20bbc7f229a19180fd3f0133e153ea828098418e260011eced0f3255a

Observation 9b258fc4-ad6b-41c4-a2b2-984d63fe9bdb · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:25.663791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:25.663791Z digest=sha256:7d5292eb8375d1ec3271916765fd784d730258a5d79e97cf10bfdd2bb0be47e8

Observation c8184ae7-d659-4ff8-9a4a-a3a514c14322 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.043454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:76215d9b8d7dd0c117e5e24989f87b6ebcd9d5cbb97fe960beb45cd5d471137b

Observation d27ba4df-1089-46e6-9ff2-471ec59aa2f9 · inbound

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs cites this paper.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.914962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:14c548542667f64c8d8d100a9b2d20f9beeb937b7b984ee57797811bc633a38a

Observation 82316228-15c4-40da-99e9-3211db6570f9 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 154

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.154251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:eceaa6feb65bdbab2a440dbc53d6621228cfc4a7924dc69b3f4bf8548457dad3

Observation 2f586a10-8e87-4812-83ad-77e986131c3c · inbound

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference cites this paper.

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.483062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.483062Z digest=sha256:d74ae14b28ad939662733160c83a9578b1611c26f4877cf9ecea49a8bfffcdee