Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:55:07.904113Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2505.08620.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:55:07.904113Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f69522c5-a6ee-43d5-a9a1-09703cf97739 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b67e10b3-03eb-47d7-b597-d237c942fcbe · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f02cb7-83f1-4776-a5a3-19333a2e918e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b044020c-f4e1-4f5b-acae-1e4b273b55ca · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a019d1c3-53cb-437e-b95b-39da7b8f67a9 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 390bab73-f08d-41e8-8a53-0050a6a1e801 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c448fa03-738e-4332-b0ff-a203d72bc99f · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b612e36d-a5c9-4395-b7d3-b36f369f8b7e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241426de-e55b-4c35-ac6e-f4c579804349 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed282d8c-f2bb-4639-a96f-a74561dd6492 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference DB-LLM: Accurate Dual-Binarization for Efficient LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1032a04-a56d-4a91-8e36-13829a648ead · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Exploring Quantization for Efficient Pre-Training of Transformer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afc2dce-7266-4d10-a6c4-0d7f140e64d3 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Training deep neural networks with low precision multiplications
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9f9111-fc94-4566-a901-0eef4b546c9a · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 317cd469-d397-4b2f-a865-7a915706a715 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72338d91-ff4f-4743-a570-8e04237cdb1c · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Extreme Compression of Large Language Models via Additive Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e481b8b8-95d1-49e2-94c4-40643a0cb9a8 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Training with Quantization Noise for Extreme Model Compression
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9995e763-ac56-4686-a063-0205d52a322e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ee1dbec1-dc4e-460a-9dbd-cd5d45772b9b · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c1d871-38e5-4d10-aba5-a23e6bff121e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Geman and G
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539cfbdb-cf9a-4669-81d2-011086da64dd · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Gemma 2: Improving Open Language Models at a Practical Size
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0a92c0-193d-4bd1-8410-948fb423db71 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b255e50-2ccd-4644-b7e1-772928d9ca39 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bee96acd-57cb-44b9-bc1d-723ba9b8e24d · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf4727b-b398-4cd7-88de-8f469956f3cf · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2492e1ea-9beb-4a3c-b2f4-52aa7ac82601 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c14984f-ced2-4bee-9d0a-e222d17ffae7 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference OLMo: Accelerating the Science of Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6628a72e-73b2-43af-b950-afb68f0b51f6 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ef0392-1fcc-4b5c-b08e-647a2b76e74b · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Efficiently Modeling Long Sequences with Structured State Spaces
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75d3048-5815-4775-b017-ed02bbb6d2fb · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Textbooks Are All You Need
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbd6b45-3fa2-45a3-8dd1-d4e5c61d9605 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 13505d0a-3213-4536-a225-9a1d7d10215f · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b67e2d-1fec-48bd-a0b9-410f2b5a01c3 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72520311-25a9-4c86-82e3-a70da4991417 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b512627-f919-4499-bffc-84527ac9d3e2 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31b60467-2f1c-48dc-8bba-b80e586c3443 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5637d4-01c8-4b55-8dd3-a4921b77e2d9 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2e26410-c174-4e2f-9a8f-48983be6e020 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0581e8b3-7625-4085-869d-d0bbda431e7c · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8030413c-c3af-4266-b404-9671b20a4ef9 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Mistral 7B
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91623583-dc4e-47ec-b501-ead0a05400bf · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference A Comprehensive Evaluation of Quantization Strategies for Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84fc88d0-9a2a-44a7-98c2-51c5908f518c · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Scaling Laws for Neural Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a22e979f-9033-4976-88f9-88dd4860720f · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference BERT Busters: Outlier Dimensions that Disrupt Transformers
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd97295-052b-48ee-a1d6-9311a89c0d33 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88b5e1e-c90d-4b8c-8475-92eaa9dec5a1 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Continuous Approximations for Improving Quantization Aware Training of LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1c95a6d2-a684-4f1b-aef2-4a70c054fa29 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f4fbe98-243a-4f87-99c3-d3d6773e5afe · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f8c64e-cd0c-44b1-a2b5-888e4359c5a2 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7236cffe-940f-406e-ad5c-63f2ecc2844d · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference DeepSeek-V3 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d94162-12e3-4817-ba77-2e12764f43ec · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba550af-65c3-4b9f-aab6-7ef3e791b268 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683e553e-1eea-4b30-a8e2-392558263909 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference FP8 Formats for Deep Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae36f70-9e16-48f6-a26d-3f7e5633bdec · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference A White Paper on Neural Network Quantization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf99f63-2438-4240-a0f0-29eddb7c528b · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference 2 OLMo 2 Furious
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b743a7-f7db-4c10-88b1-798fc318679e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference RWKV: Reinventing RNNs for the Transformer Era
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5b0c0e-2db7-4bbd-8de5-f6fc9c49b100 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference FP8-LM: Training FP8 Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afee1ad-9f2f-47ff-bcb8-4c22d24e1c22 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7a0049-4f22-4e93-bf9c-61ee84c54c6e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference PB-LLM: Partially Binarized Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddeb10b5-e3e7-4d32-97f1-1b75c32f4ca7 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec3369c-b2ad-43fd-b81a-0d9204c7ad44 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df39d202-0046-41ee-887b-6abd5c9aa80d · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference LLaMA: Open and Efficient Foundation Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fff1632-0f20-4d0a-b451-838739deef44 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a72d16-bfc1-49c1-a365-f495bba7c8a8 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference FP8 versus INT8 for efficient deep learning inference
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c743b5c6-75ca-4705-b5a7-269323e77e4d · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb280529-c917-4e18-9ceb-8f41d2044b19 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21cfa1a-5be7-4884-ba60-aff6a251ca15 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference BitNet: Scaling 1-bit Transformers for Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402abf5c-5084-427c-a65a-d9f20c138853 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Emergent Abilities of Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31570def-bd1d-4203-a5d6-ef675b01d94d · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b24241-8fda-405e-8762-3237e277b414 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f23f9d0f-c4b7-4dc2-b8ac-ddcaf324ea13 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0952677b-5ac0-45bf-bdc4-34e0726f8c96 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ec0f7b-73cb-4e47-bc37-5d2b3f5d44d3 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b31ac0-1b8e-471c-a3b7-6cd45ee75356 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9275a2-7580-49fe-a2fa-754f1f608c9b · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f74d1cf-aad8-4c85-9702-28eb2d240c09 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988529cf-de0d-4dc4-b941-e1875361a074 · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference OPT: Open Pre-trained Transformer Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c61e00b-4e12-4b9c-ae6b-80370319db2e · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference A Survey on Efficient Inference for Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f175c5c-0b68-40bc-82a2-495a124fb0cb · outbound
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference A Survey on Model Compression for Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.