Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:30.919036Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.03867.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:30.919036Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9c12499-2e2b-41b5-a52a-a149e625ea0b · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AMD Instinct™ MI350 Series GPUs,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 38f6080c-a363-4330-adc0-61b81412f0f8 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference System Card: Claude Opus 4 & Claude Sonnet 4,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 83a2abac-bb0c-4fd5-a249-0cb396aaaee0 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference QuaRot: outlier-free 4-bit inference in rotated llms,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 949f57d1-a545-4f46-8424-af02b4034e2e · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Cacti 7: New tools for interconnect exploration in innovative off-chip memories,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98bff83-0c8f-406b-ba1f-0fbe50a22354 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Piqa: Reasoning about physical commonsense in natural language,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8fc5a00d-0a28-4996-9cd6-cc39ec0b7450 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Int v.s. fp: A comprehensive study of fine-grained low-bit quantization formats,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2c8913-fb8f-4491-af78-b5ec33739032 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48277872-e5ae-49e4-8e7e-4aab138abb8c · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a7d5da-ac93-47c1-997c-294233991dce · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83d7654-d208-40e5-858c-cd7b06bc6112 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Efficient precision-scalable hardware for microscaling (mx) processing in robotics learning,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0573c81-314f-4b28-a94e-022414ad12b4 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference With shared microexponents, a little shifting goes a long way,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479b6f3c-623a-4d54-bc5a-2de962b1b321 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Deepseek-v4: Towards highly efficient million-token context intelligence,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b762538d-8fec-4633-999c-c990af96b4c5 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference LLM.int8(): 8-bit matrix multiplication for transformers at scale,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d52327e-a994-4712-a432-e8aace3f64cc · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b83e9f0c-927d-4c3f-8660-7c00c0204d1c · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Extreme compression of large language models via additive quantization,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35899f3b-85b2-4c86-97ea-076263ccede1 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf638634-a821-4b27-9813-0eabcd400783 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The language model evaluation harness,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8c3788-d6d7-416c-a0aa-37bb6bd0f321 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 4 Model Overview,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e7b05792-47c7-4ea6-b42d-3ec3fdfafa03 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ANT: Exploiting adaptive numerical data type for low-bit deep neural network quantization,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8029f34d-20a4-4469-bd82-df0ecdd958b5 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference BBAL: A bidirectional block floating point-based quantisation accelerator for large language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125c4c63-2891-4ec1-a961-9cf9aef6f499 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Measuring Massive Multitask Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f515fe-86d0-4f6a-870e-042bcc1a7de5 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 1.1 computing’s energy problem (and what we can do about it),
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17875c65-6793-40a3-bcbd-cb21bf5e4199 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4770c4e1-84c9-40a2-820f-f86281e27bae · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M-ANT: Efficient low-bit group quantization for llms via mathematically adaptive numerical type,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8a9df075-b593-4db8-b489-565517638e7b · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference M2XFP: A metadata-augmented microscaling data format for efficient low-bit quantization,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1f572b-6924-49b8-b55a-265e253bf2fd · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Blockdialect: block-wise fine-grained mixed format quantization for energy-efficient llm inference,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1942b81a-3f95-47fc-ac60-9edbe07e590e · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A diagram is worth a dozen images,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5fb49d37-4664-4019-8491-513e06c21bb3 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 14.2 a 16nm 216kb, 188.4tops/w and 133.5tflops/w microscaling multi- mode gain-cell cim macro edge-ai devices,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9293ecb2-7281-40f4-9ea4-f69183e2d117 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SqueezeLLM: dense-and-sparse quantiza- tion,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1f446c27-ee3b-44c2-b0aa-87ce14dce278 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Tender: Accelerating large language models via tensor decomposition and runtime requantization,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 68b802df-f590-4a37-8d47-e15787d70166 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference MX+: Pushing the limits of microscaling formats for efficient large language model serving,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b42c4bc-c813-4322-9aeb-87e3d7147472 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference AWQ: Activation-aware weight quantization for on-device llm compression and acceleration,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d175522-191d-485b-864f-3e336d846344 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe303a0a-7967-48ef-bba0-6916b10b4dd8 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da2dd973-057b-4547-84f8-d408236063a4 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf06264-3899-4c5b-ba41-699d6a484938 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pointer sentinel mixture models,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce95d21e-1406-4669-98a4-0523f4c90fca · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference The Llama 3 Herd of Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28354834-da05-47a8-a1f2-f1e269bb85d5 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Four MTIA Chips in Two Years: Scaling AI Experi- ences for Billions,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4524645d-eafa-4379-bd0b-c0b96d1facf0 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Recipes for Pre-training LLMs with MXFP8
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dd0de1-3794-4a65-81dd-6b76de128c2f · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA H100 Tensor Core GPU Datasheet,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ede43b36-648a-4bc5-8422-66a65e77d04c · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference NVIDIA Blackwell Architecture Technical Brief,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c378b2c6-7f03-425e-8096-0634641038ae · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OCP Microscaling Formats (MX) Specification,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 96f60da0-4c40-4a88-a26d-c01fe9b2c159 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OpenAI GPT-5 System Card,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9235f464-15d1-4964-82b4-eab6716076e7 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Paszke, S
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 026e5800-ad1d-4268-b9f8-e02e0ec0b35a · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Scale-sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6faf28f5-d8e9-49bf-8d03-f2d1878ab052 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscopiq: Accelerating foundational models through outlier-aware microscaling quantization,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 988dd783-6710-4168-9800-8fa23fba956c · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Gemma 2: Improving Open Language Models at a Practical Size
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d5b7539-4204-4fde-a6af-5cf1f3177c2b · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f8149f9c-a0c1-4b1e-84be-2c3e9c5b26d4 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Microscaling Data Formats for Deep Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b68ba9-8e7e-4d30-9af9-23f8051ab664 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Winogrande: an adversarial winograd schema challenge at scale,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c3f1d7f-18c7-44ab-b77e-62787b4c2eaa · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c817380b-ca78-47b7-a649-d41cb3bb13ef · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Towards vqa models that can read,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea0eed34-34b1-4b9b-b1c9-15b74de6e25a · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Commonsenseqa: A question answering challenge targeting commonsense knowledge,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92f0e344-78fd-40aa-ac6a-384dabf9c983 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference A microscaling multi-mode gain-cell computing-in- memory macro for advanced ai edge device,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0f9c46a0-8a74-49b2-ad29-c854ada80fd8 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8c0dbc9-5fda-4415-a49b-b34e1914798e · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference 30.1 a 28nm 127.54tflops/w mxfp6 and 117.42tflops/w mxfp8 compute-in-memory macro with adaptive- preserved-bit-width and serial-dual-bit-sliding schemes,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e41a0196-9109-443b-97d9-14fa97783159 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Accelergy: An architecture- level energy estimation methodology for accelerator designs,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e2bea2-e0a6-42cc-b28d-a2181327298e · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 14e1b5ac-698f-466e-81f7-f6a81c81dbbc · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Inside Maia 100: Revolutionizing AI Workloads with Microsoft’s Custom AI Accelerator,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce701b2b-ff15-4e2e-ac69-3e6a1a2eefa3 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference An empirical study of microscaling formats for low-precision llm training,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a94af574-4fab-41b3-be30-29247e664c8b · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Qwen2.5 Technical Report
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4480135-ffae-484b-a485-d5b94840eeef · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference ZeroQuant: efficient and affordable post-training quantization for large- scale transformers,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c371b86c-9450-4eaa-aa06-56caf514d875 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a5732eb2-2ab1-4f99-a65c-9ab9668675d8 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Hellaswag: Can a machine really finish your sentence?
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2fedea90-72f7-44d1-b563-6394105f8d4b · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0425dbc-440e-4ee4-84bc-020d68b61a3d · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Atom: Low-bit quantization for efficient and accurate llm serving,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ebe5e4cd-69d1-4586-a98f-f71b09d188b4 · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2a2b2c-93a8-4238-84da-dbd2ecbd8d5d · outbound
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Available: https://resources.nvidia.com/en-us-gpu- resources/h100-datasheet-24306
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.