Pith. sign in

Paper Citation Record · LEDGER

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition

As of 16 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2505.08981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08981 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:23.218413Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbd54209-21f3-4fed-9405-838acd6c5b68 · outbound

This paper cites Rae, Oriol Vinyals, and Laurent Sifre.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Rae, Oriol Vinyals, and Laurent Sifre

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.118998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.118998Z digest=sha256:4f3cc1aab7e76387bc22e8d829056d760186556eddb2d68df81525a8d032a8a6

Observation c6beb585-68ae-4fd4-a4b5-558043069dde · outbound

This paper cites M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.564892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.124499Z digest=sha256:a46ab6a95b4992fac888683608fc45d6b8da4addd6d5f515778c8c8184134dbc

Observation b1d16bfc-9ecb-420f-b000-ed94bec445ba · outbound

This paper cites Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.549970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.129585Z digest=sha256:65261e1d24c9d663aaf9db047afb00a4c2bc2adb9a43aee6832246d029aa6332

Observation dee71ee8-47f8-45c7-afa4-e18926d0dec5 · outbound

This paper cites Democratizing neural ma- chine translation with OPUS-MT.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Democratizing neural ma- chine translation with OPUS-MT

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.534855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.134568Z digest=sha256:82363ebe19ff92ac49e485252120a389de1df3c588646a741f9e9a204043eca1

Observation ad9b9f58-fb11-4ec6-9f3c-1c534041a094 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Omniquant: Omnidirectionally calibrated quantization for large language models, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.139139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.139139Z digest=sha256:2acd8ac9bc94269e593fb9e1eb76f1b43c7a947e48c4a7fc5f7e2e0de585e8df

Observation b3fff66a-3173-4584-9548-a39e5e29db75 · outbound

This paper cites Qllm: Accurate and efficient low-bitwidth quantization for large language models, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Qllm: Accurate and efficient low-bitwidth quantization for large language models, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.506708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.144367Z digest=sha256:8d92d2d9e39fba488b6a415c35d1ed2b7a762ee4a2b04745819e7772ce50411e

Observation 6f3fc791-5796-4577-a26e-910f529db4d2 · outbound

This paper cites Efficient arbitrary precision acceleration for large language models on gpu tensor cores, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Efficient arbitrary precision acceleration for large language models on gpu tensor cores, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.489205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.149667Z digest=sha256:d240f0dc1109f49b9b5fb688f2c2a79372216c930f8e70c64a6316e68777c18f

Observation e4898226-3c6f-4410-be4d-42511c865935 · outbound

This paper cites Q8bert: Quantized 8bit bert.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q8bert: Quantized 8bit bert

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.474136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.154624Z digest=sha256:e4795d5dc4a4b043cf1d638a72a758dd7465cd8406eeca2aa0a97527af86fa86

Observation aac596a7-35fc-41dd-b2d3-efa5553c8959 · outbound

This paper cites Q-bert: Hessian based ultra low precision quantization of bert.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q-bert: Hessian based ultra low precision quantization of bert

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.458192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.159208Z digest=sha256:7376f48911ed1d9feed9d6e88429cd72aa4d21faded341bc1472f5ff2665c698

Observation 09eb866a-4af7-47ea-b2e4-40e00e3fa77b · outbound

This paper cites Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.443008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.163729Z digest=sha256:449e22b9072586d1f2117ee8d4c2fac801ebbc57d68e9d16516540d6e897cf2a

Observation 96c12863-bda7-490f-a7ff-1c3cdec005a2 · outbound

This paper cites Hawq: Hessian aware quantization of neural networks with mixed-precision.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hawq: Hessian aware quantization of neural networks with mixed-precision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.168431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.168431Z digest=sha256:7aa77a5b17e4d9fc31ba43fa673a5390168a4b80b9153f2e33793b66f228d228

Observation f4ac93e7-a401-4204-8c39-840f13031454 · outbound

This paper cites BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:46:23.285813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.173339Z digest=sha256:63a06daf0a4d6e2fdc1ab0fd03db2a6fa6aa28319e0dce658935b097743b365a

Observation 51426ebb-e62d-4c45-9137-6b920764db9b · outbound

This paper cites Optimizing bit-serial matrix multiplica- tion for reconfigurable computing.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Optimizing bit-serial matrix multiplica- tion for reconfigurable computing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.415450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.178681Z digest=sha256:6653da6e72269a8eb73d50029266fe1d983753b98290422d2db1c7b5e4ec38d9

Observation 60f872e6-3a1f-403a-a971-02abb297cd71 · outbound

This paper cites Hihispmv: Sparse matrix vector multiplication with hierarchical row reductions on fpgas with high bandwidth memory.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hihispmv: Sparse matrix vector multiplication with hierarchical row reductions on fpgas with high bandwidth memory

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.399802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.183581Z digest=sha256:96ea78d08710da1a13696fa7225041b2e54d22f484ad9951a4886963f9df5f47

Observation ad0b3b13-9f5f-4fd5-97c6-07321e878da9 · outbound

This paper cites HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:46:23.263510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.188585Z digest=sha256:82b8c339fedf67e4ad8c1f224591038a3f74cf6ca5ca59b51e9ff581abde4511

Observation 90de80ed-2264-47ba-883d-ce89c42b9af4 · outbound

This paper cites Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.384463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.193857Z digest=sha256:52096c3b377fcd197e5fe1de588ff54ae8ac863376918a74b38c90f74e79e67c

Observation 275824aa-a5a4-4876-a287-29f5aae853ee · outbound

This paper cites Streamsvd: Low-rank ap- proximation and streaming accelerator co-design.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Streamsvd: Low-rank ap- proximation and streaming accelerator co-design

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.366096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.200030Z digest=sha256:e2d524f27f4bc3462f45e6b2233c271f359f474a64f9af4d7030608672713ba3

Observation f5ba8875-fc91-4098-973f-55f9f27a509d · outbound

This paper cites Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architec- ture, 2023.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architec- ture, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.349573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.204691Z digest=sha256:8a9564cd2ba7609c31980955bee3351bbd4d0e394f5a8b69b85f3177bec99b15

Observation 2473c3b1-39c4-4243-b2b9-d609f08395be · outbound

This paper cites Film-qnn: Efficient fpga acceleration of deep neural networks with intra-layer, mixed-precision quantization.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Film-qnn: Efficient fpga acceleration of deep neural networks with intra-layer, mixed-precision quantization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.332969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.209081Z digest=sha256:796b34858e10b7edda9bbb2de46cb83c2e7258201e6aa343dae8333144f9e5cb

Observation 3ebbecc7-6dbf-4877-ae6d-e4c0ae4d8643 · outbound

This paper cites Understanding the potential of fpga-based spatial acceleration for large language model inference.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Understanding the potential of fpga-based spatial acceleration for large language model inference

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.317160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.213840Z digest=sha256:95d6adbe6b9cfdaecca40b0f92e9351b63a9ae25755717403077e53107e598e9

Observation 1fe1fd53-c6bd-44b1-a568-66745d185d29 · outbound

This paper cites an unresolved cited work.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:23.301735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:46:23.218413Z digest=sha256:fd9022c86717b2e43874cbee195537f1b02549270719a35d44322787a6b6d1a1

Pith citing papers

No inbound Pith citation observations are available.