Pith. sign in

Paper Citation Record · LEDGER

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2505.08981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08981 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:23.218413Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbd54209-21f3-4fed-9405-838acd6c5b68 · outbound

This paper cites Rae, Oriol Vinyals, and Laurent Sifre.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Rae, Oriol Vinyals, and Laurent Sifre

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.118998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.118998Z digest=sha256:4f3cc1aab7e76387bc22e8d829056d760186556eddb2d68df81525a8d032a8a6

Observation c6beb585-68ae-4fd4-a4b5-558043069dde · outbound

This paper cites M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.564892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.124499Z digest=sha256:5c22facd517be075c69b6e5d828b22f86535dfc8a4ada53d2c4a2870ee8937f5

Observation b1d16bfc-9ecb-420f-b000-ed94bec445ba · outbound

This paper cites Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.549970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.129585Z digest=sha256:554f50f708d6fb96c404010e577c4b0ce9e3bf5223b0b0b7f39c38a3785e9f29

Observation dee71ee8-47f8-45c7-afa4-e18926d0dec5 · outbound

This paper cites Democratizing neural ma- chine translation with OPUS-MT.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Democratizing neural ma- chine translation with OPUS-MT

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.534855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.134568Z digest=sha256:9f6607a408ddeb6ab3558646481cdf948ad0d704cc6d865669f24bb58a628d79

Observation ad9b9f58-fb11-4ec6-9f3c-1c534041a094 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Omniquant: Omnidirectionally calibrated quantization for large language models, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.139139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.139139Z digest=sha256:2acd8ac9bc94269e593fb9e1eb76f1b43c7a947e48c4a7fc5f7e2e0de585e8df

Observation b3fff66a-3173-4584-9548-a39e5e29db75 · outbound

This paper cites Qllm: Accurate and efficient low-bitwidth quantization for large language models, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Qllm: Accurate and efficient low-bitwidth quantization for large language models, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.506708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.144367Z digest=sha256:7143801304eb28d26628862a64a1db03c241ab020286c0c5740e31cec6b2dc1b

Observation 6f3fc791-5796-4577-a26e-910f529db4d2 · outbound

This paper cites Efficient arbitrary precision acceleration for large language models on gpu tensor cores, 2024.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Efficient arbitrary precision acceleration for large language models on gpu tensor cores, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.489205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.149667Z digest=sha256:078987aa703323cd063e29ab164dfa3ddd4b152c9ade6dd2984ec857b30fce5a

Observation e4898226-3c6f-4410-be4d-42511c865935 · outbound

This paper cites Q8bert: Quantized 8bit bert.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q8bert: Quantized 8bit bert

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.474136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.154624Z digest=sha256:4218d044fa00f37870aa638c62f125216fb7317ce39e9d494dc05ad2212d64f2

Observation aac596a7-35fc-41dd-b2d3-efa5553c8959 · outbound

This paper cites Q-bert: Hessian based ultra low precision quantization of bert.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Q-bert: Hessian based ultra low precision quantization of bert

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.458192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.159208Z digest=sha256:fbbb8f84809b777358e8482ad08bdbc089d0352017f90ef6b033ed26da049c99

Observation 09eb866a-4af7-47ea-b2e4-40e00e3fa77b · outbound

This paper cites Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.443008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.163729Z digest=sha256:4132b38c14051af0a71139aae053d906c9d1d06f6576c6865318388a9357cc3b

Observation 96c12863-bda7-490f-a7ff-1c3cdec005a2 · outbound

This paper cites Hawq: Hessian aware quantization of neural networks with mixed-precision.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hawq: Hessian aware quantization of neural networks with mixed-precision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:23.168431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:23.168431Z digest=sha256:7aa77a5b17e4d9fc31ba43fa673a5390168a4b80b9153f2e33793b66f228d228

Observation f4ac93e7-a401-4204-8c39-840f13031454 · outbound

This paper cites BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:46:23.285813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.173339Z digest=sha256:d5654dd2659348f49a1e573d30b4d5176aa9390e2a30891b267a3149bf8747f3

Observation 51426ebb-e62d-4c45-9137-6b920764db9b · outbound

This paper cites Optimizing bit-serial matrix multiplica- tion for reconfigurable computing.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Optimizing bit-serial matrix multiplica- tion for reconfigurable computing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.415450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.178681Z digest=sha256:671076cf404a6639b30416a992efcf2d2289f185a7485d28ca8b072bcd0042c0

Observation 60f872e6-3a1f-403a-a971-02abb297cd71 · outbound

This paper cites Hihispmv: Sparse matrix vector multiplication with hierarchical row reductions on fpgas with high bandwidth memory.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Hihispmv: Sparse matrix vector multiplication with hierarchical row reductions on fpgas with high bandwidth memory

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.399802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.183581Z digest=sha256:d7439c186e773f6128a91574da84c2fe5e072401399cd9f05a78e91ac4d5fa00

Observation ad0b3b13-9f5f-4fd5-97c6-07321e878da9 · outbound

This paper cites HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:46:23.263510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.188585Z digest=sha256:e3733f98b717d79caa4c13760acf79aa8eb8ab9312a89fcdb00c3210658adb25

Observation 90de80ed-2264-47ba-883d-ce89c42b9af4 · outbound

This paper cites Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.384463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.193857Z digest=sha256:6d32389d45b7da6f3e75f96dc878162c783b10ee358ac4b1636992279387d10c

Observation 275824aa-a5a4-4876-a287-29f5aae853ee · outbound

This paper cites Streamsvd: Low-rank ap- proximation and streaming accelerator co-design.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Streamsvd: Low-rank ap- proximation and streaming accelerator co-design

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.366096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.200030Z digest=sha256:da4e9fa19bc01a81da495939b5439c819aa09421e1931cce31c709bd82046119

Observation f5ba8875-fc91-4098-973f-55f9f27a509d · outbound

This paper cites Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architec- ture, 2023.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architec- ture, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.349573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.204691Z digest=sha256:c25d5875f6ca2a18a08f382a7a2767eea2b8ed462889482d4a22cf4b7bf761ad

Observation 2473c3b1-39c4-4243-b2b9-d609f08395be · outbound

This paper cites Film-qnn: Efficient fpga acceleration of deep neural networks with intra-layer, mixed-precision quantization.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Film-qnn: Efficient fpga acceleration of deep neural networks with intra-layer, mixed-precision quantization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.332969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.209081Z digest=sha256:ea847b16f3217caa3a18378eba8176dfb6f1e403349bdb868c9d5c779f824c53

Observation 3ebbecc7-6dbf-4877-ae6d-e4c0ae4d8643 · outbound

This paper cites Understanding the potential of fpga-based spatial acceleration for large language model inference.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Understanding the potential of fpga-based spatial acceleration for large language model inference

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:23.317160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.213840Z digest=sha256:5446a82c937e493f5bfabea612442222cd9f17c5701c79147a8b973b0cda7132

Observation 1fe1fd53-c6bd-44b1-a568-66745d185d29 · outbound

This paper cites an unresolved cited work.

ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:23.301735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:46:23.218413Z digest=sha256:8126450e7cfd037db6639022e241dee96b277038f4acdf54ce44ec33ccfa83d9

Pith citing papers

No inbound Pith citation observations are available.