Pith. sign in

Paper Citation Record · LEDGER

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.03867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03867 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:30.919036Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9c12499-2e2b-41b5-a52a-a149e625ea0b · outbound

This paper cites AMD Instinct™ MI350 Series GPUs,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AMD Instinct™ MI350 Series GPUs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.254496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:28.061738Z digest=sha256:bba5a703032a0ffe94fce769da1255d192bcf8b9dfcc388f711ba84b80c7deb8

Observation 38f6080c-a363-4330-adc0-61b81412f0f8 · outbound

This paper cites System Card: Claude Opus 4 & Claude Sonnet 4,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference System Card: Claude Opus 4 & Claude Sonnet 4,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.245420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:28.148840Z digest=sha256:433a5cd9a0323930d737664a72b37d84d1221643bbc54284822f75c946c2fc78

Observation 83a2abac-bb0c-4fd5-a249-0cb396aaaee0 · outbound

This paper cites QuaRot: outlier-free 4-bit inference in rotated llms,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference QuaRot: outlier-free 4-bit inference in rotated llms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.236474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:28.229286Z digest=sha256:327b0c7021165d2eeba1bd038cad9b2e20bcc18eb82f32d00857c5439532f9ac

Observation 949f57d1-a545-4f46-8424-af02b4034e2e · outbound

This paper cites Cacti 7: New tools for interconnect exploration in innovative off-chip memories,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Cacti 7: New tools for interconnect exploration in innovative off-chip memories,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.263418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.263418Z digest=sha256:9efd06bd40f983c4ffb80ad5254ee24f64a076eda593886b70d1c125c891ae53

Observation d98bff83-0c8f-406b-ba1f-0fbe50a22354 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Piqa: Reasoning about physical commonsense in natural language,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.227307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:28.329216Z digest=sha256:2f06665dd8b43cebb0da3b9512aec4b5d3e285c841442784b2a5e01a014c25cd

Observation 8fc5a00d-0a28-4996-9cd6-cc39ec0b7450 · outbound

This paper cites Int v.s. fp: A comprehensive study of fine-grained low-bit quantization formats,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Int v.s. fp: A comprehensive study of fine-grained low-bit quantization formats,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.401800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.401800Z digest=sha256:8f5a775017d844a5ee5ef5bb4f5ddde608491aa0cfbbbe7ac07b55e607aa2674

Observation 5c2c8913-fb8f-4491-af78-b5ec33739032 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.471068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.471068Z digest=sha256:9a9871b738d8cf4cfc3acb71b5d43fa2a4f09b2a60ebc963f69b1c178728b548

Observation 48277872-e5ae-49e4-8e7e-4aab138abb8c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.580769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.580769Z digest=sha256:5a5b6d9e41488bd3b05361cec54f2db0b0578faf2aedaf9f60a11571ee4c27eb

Observation f9a7d5da-ac93-47c1-997c-294233991dce · outbound

This paper cites Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.676839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.676839Z digest=sha256:82d69c4c8ddc327ad2415226c7feb836fb45b881e0617f5c38c972627d0b9f7c

Observation a83d7654-d208-40e5-858c-cd7b06bc6112 · outbound

This paper cites Efficient precision-scalable hardware for microscaling (mx) processing in robotics learning,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Efficient precision-scalable hardware for microscaling (mx) processing in robotics learning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.699322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.699322Z digest=sha256:ce6a602f99ff34b76fbeb6a38eeedca67d3a4aa549eed014fe689ae118721e50

Observation d0573c81-314f-4b28-a94e-022414ad12b4 · outbound

This paper cites With shared microexponents, a little shifting goes a long way,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference With shared microexponents, a little shifting goes a long way,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.792184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.792184Z digest=sha256:ec3585c2a4ee97dcbd3ab472ed8ca27bd1e85272a524a47f262b997235809f21

Observation 479b6f3c-623a-4d54-bc5a-2de962b1b321 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Deepseek-v4: Towards highly efficient million-token context intelligence,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:28.907764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:28.907764Z digest=sha256:0b6888d2b67805ac068576919b995d9aa4aa8fa8e90df86cabfda554a6a15495

Observation b762538d-8fec-4633-999c-c990af96b4c5 · outbound

This paper cites LLM.int8(): 8-bit matrix multiplication for transformers at scale,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference LLM.int8(): 8-bit matrix multiplication for transformers at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.212025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:29.005235Z digest=sha256:f102a000199225e0b1c0906329d5ffef95da488f1e771ad7b6b446e098a57c65

Observation 0d52327e-a994-4712-a432-e8aace3f64cc · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.065972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.065972Z digest=sha256:bf1fb1aa93d65c19cc141668e57fae0faa666fb8ab978c34a70e8f7c2af4ea75

Observation b83e9f0c-927d-4c3f-8660-7c00c0204d1c · outbound

This paper cites Extreme compression of large language models via additive quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Extreme compression of large language models via additive quantization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.202728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:29.143132Z digest=sha256:78186478eaf70cbc9acdcae698cd9e1a45f3448403987b9e74e709f9992f99c6

Observation 35899f3b-85b2-4c86-97ea-076263ccede1 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.303924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.303924Z digest=sha256:adbcde15924c2033a70d9ced2553a479fc471d2957f427e9cb61364c91205a1e

Observation cf638634-a821-4b27-9813-0eabcd400783 · outbound

This paper cites The language model evaluation harness,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The language model evaluation harness,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.417903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.417903Z digest=sha256:fe127917ed338b174006d95abe2112c3149f6e664aed2b463bfd8f8e1728d9ce

Observation 2a8c3788-d6d7-416c-a0aa-37bb6bd0f321 · outbound

This paper cites Gemma 4 Model Overview,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 4 Model Overview,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.193835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:29.550359Z digest=sha256:4f397f82d7bc9febafbe18fa68b66d04bb2fd25a8be3c9f28a6fa890905a4a61

Observation e7b05792-47c7-4ea6-b42d-3ec3fdfafa03 · outbound

This paper cites ANT: Exploiting adaptive numerical data type for low-bit deep neural network quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ANT: Exploiting adaptive numerical data type for low-bit deep neural network quantization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.184673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:29.637719Z digest=sha256:35d6a99fe628498aff276d9962fc1997b4cf52474fceea9d3cc7c51f52d55f27

Observation 8029f34d-20a4-4469-bd82-df0ecdd958b5 · outbound

This paper cites BBAL: A bidirectional block floating point-based quantisation accelerator for large language models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BBAL: A bidirectional block floating point-based quantisation accelerator for large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.770682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.770682Z digest=sha256:54f8635311ee1b3babe758c09b4470fd4271e257f8a30d6959a243785c4e846d

Observation 125c4c63-2891-4ec1-a961-9cf9aef6f499 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:29.896018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:29.896018Z digest=sha256:241bdafdafd7c57f666157da47e7231d53a7b7626a5bd1cdc728eb17878f9bc9

Observation d5f515fe-86d0-4f6a-870e-042bcc1a7de5 · outbound

This paper cites 1.1 computing’s energy problem (and what we can do about it),.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 1.1 computing’s energy problem (and what we can do about it),

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.016464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.016464Z digest=sha256:f565dd364ded45074e9eddb5e0ca249a79c32831e1fa1a927ad0f8c8bbbbf7b7

Observation 17875c65-6793-40a3-bcbd-cb21bf5e4199 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.113712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.113712Z digest=sha256:5c0fbc684976fc0f0ddc8c62ae2a9183f0abad6560e7d16e25ef4f5bd576455c

Observation 4770c4e1-84c9-40a2-820f-f86281e27bae · outbound

This paper cites M-ANT: Efficient low-bit group quantization for llms via mathematically adaptive numerical type,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M-ANT: Efficient low-bit group quantization for llms via mathematically adaptive numerical type,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.169110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.264778Z digest=sha256:59d909b930f3029cf6b5a981a42fd20c685ecc6d89880b079b560393068f0e4f

Observation 8a9df075-b593-4db8-b489-565517638e7b · outbound

This paper cites M2XFP: A metadata-augmented microscaling data format for efficient low-bit quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M2XFP: A metadata-augmented microscaling data format for efficient low-bit quantization,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.375670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.375670Z digest=sha256:490cfd0bb5c789f4a65b664cd4898dcbe980bac665e93498e02ccd94b9d91fcf

Observation 3c1f572b-6924-49b8-b55a-265e253bf2fd · outbound

This paper cites Blockdialect: block-wise fine-grained mixed format quantization for energy-efficient llm inference,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Blockdialect: block-wise fine-grained mixed format quantization for energy-efficient llm inference,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.158601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.535766Z digest=sha256:d2cf05d8dcfa5dfc0dd41beef67417b4d63581f7833724ed6bd24e34dd092845

Observation 1942b81a-3f95-47fc-ac60-9edbe07e590e · outbound

This paper cites A diagram is worth a dozen images,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A diagram is worth a dozen images,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.147849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.700003Z digest=sha256:7ce9366538ad5facb14856de689e4cdf86715a90f6f9c607512f15b79ff61c1f

Observation 5fb49d37-4664-4019-8491-513e06c21bb3 · outbound

This paper cites 14.2 a 16nm 216kb, 188.4tops/w and 133.5tflops/w microscaling multi- mode gain-cell cim macro edge-ai devices,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 14.2 a 16nm 216kb, 188.4tops/w and 133.5tflops/w microscaling multi- mode gain-cell cim macro edge-ai devices,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.138096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.799634Z digest=sha256:8a979af1ad0e0a956f0e6c88cd493332e9fd28957d9130c447a9e545a885c1db

Observation 9293ecb2-7281-40f4-9ea4-f69183e2d117 · outbound

This paper cites SqueezeLLM: dense-and-sparse quantiza- tion,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SqueezeLLM: dense-and-sparse quantiza- tion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.127987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.802603Z digest=sha256:31669b1faa20234444bf9fe209caba504c30c5f6169a60713511a4ba986ed595

Observation 1f446c27-ee3b-44c2-b0aa-87ce14dce278 · outbound

This paper cites Tender: Accelerating large language models via tensor decomposition and runtime requantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Tender: Accelerating large language models via tensor decomposition and runtime requantization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.118081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.805468Z digest=sha256:9c4478354436e1f06833d75b1cbf1daefd27c53e1137caac34974868ddbd5c60

Observation 68b802df-f590-4a37-8d47-e15787d70166 · outbound

This paper cites MX+: Pushing the limits of microscaling formats for efficient large language model serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference MX+: Pushing the limits of microscaling formats for efficient large language model serving,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.808079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.808079Z digest=sha256:7f7272a2a5cee68635042a559edd1e2956e0744d46a2afca95d2f39a23025dd4

Observation 9b42c4bc-c813-4322-9aeb-87e3d7147472 · outbound

This paper cites AWQ: Activation-aware weight quantization for on-device llm compression and acceleration,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AWQ: Activation-aware weight quantization for on-device llm compression and acceleration,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.810712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.810712Z digest=sha256:7f6e97f7f89f5194db08897ab6f9da3b9c9ea454c49463d3e6b737ed3809df25

Observation 7d175522-191d-485b-864f-3e336d846344 · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.813720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.813720Z digest=sha256:9814f1ad560be6701f3db80ff2a9265596e10e17988bc6c521202a461fb4d421

Observation fe303a0a-7967-48ef-bba0-6916b10b4dd8 · outbound

This paper cites Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:54:31.243463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.816676Z digest=sha256:2b05863872e96b916331be858dcf971c398afe727095c5045b9c9e55e3d0c37c

Observation da2dd973-057b-4547-84f8-d408236063a4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.820251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.820251Z digest=sha256:46dbfe584158f8930360b2f08fddb3ad1473f1e6e1baa9e73ba9f636d4ade041

Observation dbf06264-3899-4c5b-ba41-699d6a484938 · outbound

This paper cites Pointer sentinel mixture models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pointer sentinel mixture models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.095906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.823013Z digest=sha256:6d36cd23b75aa325708cbb1164db2a080300ac546ff7d299df024ccdea62160b

Observation ce95d21e-1406-4669-98a4-0523f4c90fca · outbound

This paper cites The Llama 3 Herd of Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The Llama 3 Herd of Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.825891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.825891Z digest=sha256:4244ca2606e53e9b4b541808d8235ace0d2d674cfbab92449a37be8823826eb2

Observation 28354834-da05-47a8-a1f2-f1e269bb85d5 · outbound

This paper cites Four MTIA Chips in Two Years: Scaling AI Experi- ences for Billions,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four MTIA Chips in Two Years: Scaling AI Experi- ences for Billions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.086557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.828823Z digest=sha256:b604eef696c1b8c6a60d33b6d630f31f50d644004f3aa045c151a14039175117

Observation 4524645d-eafa-4379-bd0b-c0b96d1facf0 · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Recipes for Pre-training LLMs with MXFP8

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.832174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.832174Z digest=sha256:544f9c522770fb2919746956c2fee1837973a337001aa32724ddc1d740c82473

Observation 89dd0de1-3794-4a65-81dd-6b76de128c2f · outbound

This paper cites NVIDIA H100 Tensor Core GPU Datasheet,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA H100 Tensor Core GPU Datasheet,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.077146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.835513Z digest=sha256:ccd51bbd8a1d35d7d8a2b6027f7a1d20d64e2aaedd5253ca7c82545020412b1f

Observation ede43b36-648a-4bc5-8422-66a65e77d04c · outbound

This paper cites NVIDIA Blackwell Architecture Technical Brief,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA Blackwell Architecture Technical Brief,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.058539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.841679Z digest=sha256:fbf2d955ca6df841b75604dda09fdcc9ae54867211d1d43185734ad9a75a0a53

Observation c378b2c6-7f03-425e-8096-0634641038ae · outbound

This paper cites OCP Microscaling Formats (MX) Specification,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OCP Microscaling Formats (MX) Specification,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.048394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.844766Z digest=sha256:d54c2e8893fe9b23ba62e5ef54ac8778d5282543884c35836e7dbad4c6db5e82

Observation 96f60da0-4c40-4a88-a26d-c01fe9b2c159 · outbound

This paper cites OpenAI GPT-5 System Card,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OpenAI GPT-5 System Card,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.037525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.847687Z digest=sha256:fb1ad62fd646ca075a3783508fffdeb3d0abd5f088e36f1a158bda2bbc07be44

Observation 9235f464-15d1-4964-82b4-eab6716076e7 · outbound

This paper cites Paszke, S.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Paszke, S

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.850566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.850566Z digest=sha256:770b0fb9cc440e6a6b5f1e90b661c248489114cfb27084c1218911dab051ff6c

Observation 026e5800-ad1d-4268-b9f8-e02e0ec0b35a · outbound

This paper cites Scale-sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Scale-sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.019795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.853377Z digest=sha256:4754a5552521727272bd2fe9e63474c03e13b5f6a51a158315babafbaa02c471

Observation 6faf28f5-d8e9-49bf-8d03-f2d1878ab052 · outbound

This paper cites Microscopiq: Accelerating foundational models through outlier-aware microscaling quantization,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscopiq: Accelerating foundational models through outlier-aware microscaling quantization,

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T10:54:31.211877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.856127Z digest=sha256:6d4a4ae0d3cc187cb5c0dc8a3d7e985daa3d56d94ab106210160507a1129abba

Observation 988dd783-6710-4168-9800-8fa23fba956c · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 2: Improving Open Language Models at a Practical Size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.858903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.858903Z digest=sha256:2b53da08e18b9c2dda1efbadf44791817375e91494e1578fa80cfcba9efd8436

Observation 0d5b7539-4204-4fde-a6af-5cf1f3177c2b · outbound

This paper cites Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.008859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.862614Z digest=sha256:e79271acbe711bfe00398a4fc9069a83403ac9c4e7982235ef6880404172b069

Observation f8149f9c-a0c1-4b1e-84be-2c3e9c5b26d4 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscaling Data Formats for Deep Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.865642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.865642Z digest=sha256:87f4ad2a9e714a2913b0dfe5619b045e5bd04bd291e6b9d278636ad205a9e5f2

Observation d1b68ba9-8e7e-4d30-9af9-23f8051ab664 · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Winogrande: an adversarial winograd schema challenge at scale,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.869033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.869033Z digest=sha256:e5b9a5150176130ea1e98b65487dd667cd0d750d805e5e3560a29a5faa6a3d27

Observation 3c3f1d7f-18c7-44ab-b77e-62787b4c2eaa · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.872438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.872438Z digest=sha256:ebe4730fea96190cbdb3948b9117fe1e2a3e324afff18e3939e46cd806a405d9

Observation c817380b-ca78-47b7-a649-d41cb3bb13ef · outbound

This paper cites Towards vqa models that can read,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Towards vqa models that can read,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.998117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.876298Z digest=sha256:82d67c430526435b4e8d2cfaab27aec31f607621fe5eab645c86480a95a1686e

Observation ea0eed34-34b1-4b9b-b1c9-15b74de6e25a · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Commonsenseqa: A question answering challenge targeting commonsense knowledge,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.988088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.879465Z digest=sha256:744362387acd99011a705c667b9d0c5f196ae76249126d54149245d9983fc828

Observation 92f0e344-78fd-40aa-ac6a-384dabf9c983 · outbound

This paper cites A microscaling multi-mode gain-cell computing-in- memory macro for advanced ai edge device,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A microscaling multi-mode gain-cell computing-in- memory macro for advanced ai edge device,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.977638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.885660Z digest=sha256:0de0baa77d3b09bf92cd73caf0e75cf013874a4d4e3a0afc80792d9a7eb874a3

Observation 0f9c46a0-8a74-49b2-ad29-c854ada80fd8 · outbound

This paper cites Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.966233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.888361Z digest=sha256:a165df3b93e70bcae0b9c04d44c0f7f52e5d767f911a475df610ed0e3a7de87b

Observation c8c0dbc9-5fda-4415-a49b-b34e1914798e · outbound

This paper cites 30.1 a 28nm 127.54tflops/w mxfp6 and 117.42tflops/w mxfp8 compute-in-memory macro with adaptive- preserved-bit-width and serial-dual-bit-sliding schemes,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 30.1 a 28nm 127.54tflops/w mxfp6 and 117.42tflops/w mxfp8 compute-in-memory macro with adaptive- preserved-bit-width and serial-dual-bit-sliding schemes,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.956423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.890967Z digest=sha256:1735b94c1da4fca38a0492ac578c2b145f2bdc20ca385e92cec08f2a0eb07903

Observation e41a0196-9109-443b-97d9-14fa97783159 · outbound

This paper cites Accelergy: An architecture- level energy estimation methodology for accelerator designs,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Accelergy: An architecture- level energy estimation methodology for accelerator designs,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.894125Z digest=sha256:9ae9fe0a5a134e9058ca6be166205e3332e0455808eb8c619af54b5936d24eb7

Observation 63e2bea2-e0a6-42cc-b28d-a2181327298e · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.939793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.896961Z digest=sha256:e1859cb62a7c85dc869420aedaffe40f9ff3ecec32a4dc5fdd62ed871d61b169

Observation 14e1b5ac-698f-466e-81f7-f6a81c81dbbc · outbound

This paper cites Inside Maia 100: Revolutionizing AI Workloads with Microsoft’s Custom AI Accelerator,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Inside Maia 100: Revolutionizing AI Workloads with Microsoft’s Custom AI Accelerator,

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-05T10:54:31.096960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.899617Z digest=sha256:8759ebf97b09c4d1c17eaba9cfef152eaed7e01503ca08ccdbbcd6deebaa9c57

Observation ce701b2b-ff15-4e2e-ac69-3e6a1a2eefa3 · outbound

This paper cites An empirical study of microscaling formats for low-precision llm training,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference An empirical study of microscaling formats for low-precision llm training,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.928622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.902228Z digest=sha256:d177aa2a278c7de5877a6dab004228851cd65438baf63b560730f618fa6b16c0

Observation a94af574-4fab-41b3-be30-29247e664c8b · outbound

This paper cites Qwen2.5 Technical Report.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.904865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.904865Z digest=sha256:1c800afb13f1d2dba1e1b6a7e2fd60d2159858b34f99f9ab059ac025a7cf25f6

Observation e4480135-ffae-484b-a485-d5b94840eeef · outbound

This paper cites ZeroQuant: efficient and affordable post-training quantization for large- scale transformers,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ZeroQuant: efficient and affordable post-training quantization for large- scale transformers,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.919045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.907783Z digest=sha256:7e8ab29fff3ebccb7bba23354ce710108a2268d52c9b430ccd960f994026aa0a

Observation c371b86c-9450-4eaa-aa06-56caf514d875 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.909352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.910696Z digest=sha256:0d6368d900780f57d17df87b7e3ce29d9d19d0e2e36ffedf7eb403358938ce7c

Observation a5732eb2-2ab1-4f99-a65c-9ab9668675d8 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence?.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Hellaswag: Can a machine really finish your sentence?

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.899313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.913180Z digest=sha256:122a0ebd24ccfce15142396415faf2877d4b199f5f962ce28d0e8af5584d5a85

Observation 2fedea90-72f7-44d1-b563-6394105f8d4b · outbound

This paper cites Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.916143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.916143Z digest=sha256:a9a97958680b5591a0326ebcf57647c66a079d6ea32ddee6976c8cd8b976ee9a

Observation e0425dbc-440e-4ee4-84bc-020d68b61a3d · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:31.889954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.919036Z digest=sha256:d3ca2a730dff387c4a3ed3f3be46a7fa2401c1da6fc7ee8f21fbb4da9e8e8f3c

Observation ebe5e4cd-69d1-4586-a98f-f71b09d188b4 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.882544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.882544Z digest=sha256:f934780cd27a3038126e8422ec9df668402e0bc5df7916281fd875833c4bb5f0

Observation 2d2a2b2c-93a8-4238-84da-dbd2ecbd8d5d · outbound

This paper cites Available: https://resources.nvidia.com/en-us-gpu- resources/h100-datasheet-24306.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Available: https://resources.nvidia.com/en-us-gpu- resources/h100-datasheet-24306

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:32.068016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T10:54:30.838691Z digest=sha256:a44de249288e3a256da3028899ba111ada31b17e00ce782bf3ffd3e117ac06c1

Pith citing papers

No inbound Pith citation observations are available.