Pith. sign in

Paper Citation Record · LEDGER

I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2405.17849.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17849 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:14:46.381014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T15:05:48.315868Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f6d2e40-1ea0-4071-927c-229c79d75d7d · inbound

MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods cites this paper.

MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:14:46.381014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:14:46.381014Z digest=sha256:68fb2308424381a6f6119be4acd74cc85d349aa76283b79d62f6f68a1699afa4

Observation aaaa6f37-4691-448f-b13b-563a8c370dec · inbound

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting cites this paper.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.929890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.929890Z digest=sha256:5e80bdfbb9fba31a1462611363f11f192195a1c631b9dcbe56796da661fcbb7d

Observation 34c4d1a8-c958-497c-b69a-e086544699da · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.076730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.076730Z digest=sha256:ce32341790b3c3f43317c962851ab6e80f50f638215fb2ceaf8b671c253b3477

Observation bc52a88e-cc7f-4a45-86de-6f6024becd8b · inbound

A Survey: Towards Privacy and Security in Mobile Large Language Models cites this paper.

A Survey: Towards Privacy and Security in Mobile Large Language Models I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:16.858554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:16.858554Z digest=sha256:f7c6d0a3ac167be3ebe7431373227612b53f3064af77d5bf39fb5877e0f03e7a

Observation 490beeb3-ca10-49cc-9299-4a0cf099750c · inbound

IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference cites this paper.

IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:02:51.187405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:02:51.187405Z digest=sha256:ea825d8a26f0d8f1cf129044b7d213745ea5356f8236bb42e97d4fb3e158aa40

Observation c4cad0ed-f424-4df6-8c0a-7caca0d107ac · inbound

LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM cites this paper.

LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:49.600419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:05:11.335715Z digest=sha256:25b2735c782a3d920f3e4f2f87ab2308c9c4e5d7f484dcd0f4f6cf42d4c83f13

Observation 1689b1ce-b702-4cf5-b4d8-7aea8d58bc54 · inbound

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation cites this paper.

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:04.909793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T03:04:14.900791Z digest=sha256:145d10d34a00592a108b3f01506122d1bef2ce76262ac31f5a302ada4f995194

Observation d81ce732-3e0b-4860-864a-40cb03eb4197 · inbound

QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention cites this paper.

QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:31:13.467905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:57:13.714858Z digest=sha256:e6d8709d36e218f8d2a2169764f8382bb165b448dcc887d07794a9f27cfd3f08

Observation da198aef-495c-470c-be50-04cd1b648111 · inbound

Theory-optimal Quantization Based on Flatness cites this paper.

Theory-optimal Quantization Based on Flatness I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:09.877448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:38:06.888665Z digest=sha256:073bb834a1b995e26e1dbb2c8d64c9d1f4784e369726f3bd1783900e672f434b

Observation bd04d0b6-7d5c-4822-a995-1676b3fb7806 · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:23:03.664882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:01ebb9ee3b4965fac11b9a6413eeed471908fce0b339585e61fb50035087aee1

Observation e64e12ef-edd7-40b8-ba62-dc909f7881e7 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.318099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:2be6db1326f1e666127bc3fa1f386edd0e49bfd783202f4b9b1adb8b74a0b7f0

Observation 4ceafbf5-d55c-4ad1-958f-e7e342ded546 · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:44:45.577156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:6b359d7fff2891c7e4af8d4fbcef870eb6aea7c1ec91a4a17a417cb097cd6f07

Observation e0c1e0da-c194-42ca-8880-07fa92883e7e · inbound

LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression cites this paper.

LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T05:11:18.690792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:11:18.690792Z digest=sha256:b862b679b49fb1375de87ac6ebf6956ac4fa1f72991a190f661edcdc4520ab35

Observation c72051f7-88bf-4e14-9a50-49e59225e258 · inbound

Break Through the Compression Bottleneck: From Theory to Practice cites this paper.

Break Through the Compression Bottleneck: From Theory to Practice I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:29:28.763839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:29:28.763839Z digest=sha256:d83a6c905b701c2c627e9d2d1850a32ab844ca6e028bbc3622dc7a44fd76a387

Observation 29e8d529-e7f6-4b9e-af2f-bc81687236eb · inbound

When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation cites this paper.

When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Reference 146

Resolution
unresolved
no resolver link, observed 2026-07-30T23:38:38.615687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T23:38:38.615687Z digest=sha256:00cf87b48a35762caedaff7b07942125be2bd31fa189813bf15d7815aa764d70