Pith. sign in

Paper Citation Record · LEDGER

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.14638.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14638 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.850300Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 236c17b8-9918-47c9-aa7d-3e09aa1d7f80 · outbound

This paper cites Learned Step Size Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Learned Step Size Quantization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.491210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.491210Z digest=sha256:dc9c8818e4bfb65bb9c25174561392145461388bd770c103ffba22878bd07045

Observation 94b1badb-5fc8-4880-ac05-496904854c23 · outbound

This paper cites Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.082662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.498781Z digest=sha256:bc4e2610c376cc2c02de905f2056943a9081109f2d39577a2b1a0ac4f728df61

Observation 2c1403f1-ef5d-45e0-9004-6e41e8b4630e · outbound

This paper cites Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.056103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.506094Z digest=sha256:5d27ce8c9afd62e69165c4cda0bbf8e533341b96852469b4d37d82af872c1f67

Observation 1c5fd896-d33d-4a13-8799-4efc3d93369b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.513839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.513839Z digest=sha256:33011e9cd23bb2875490ae322b2678e850a1c36561fa629b0abaa0347a3866ba

Observation 89a56030-cdbd-4dc1-94f1-e2a0167bf1b3 · outbound

This paper cites Data-free quantization through weight equalization and bias correction,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Data-free quantization through weight equalization and bias correction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.030194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.520642Z digest=sha256:5016cfcec2b9a984d2bf182a309f98301e18262233878f79b651bb94e49b6b76

Observation e0c27305-2a9d-404e-a070-7be9039062a8 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.527071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.527071Z digest=sha256:7c3f18ecd52638d3d70cda67452b80c1b361f722918a5857274a7b54b36f83d9

Observation 13814e29-fb9e-4965-9a5a-ca9090c6622c · outbound

This paper cites Accurate post training quantization with small calibra- tion sets,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Accurate post training quantization with small calibra- tion sets,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.009418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.534639Z digest=sha256:2ad8e1c2eb89ca1458908fbe32a9b203e4d9e76fb938a0328751747ece7ff147

Observation c581bb25-a830-4af4-a641-a331f5e4e420 · outbound

This paper cites Post-training quantization for vision transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Post-training quantization for vision transformer,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.987699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.540267Z digest=sha256:0a80b7854a1bd6762cd68a59af49a4ba5c674797a65ddfd56adf8227532b4f04

Observation 0475f67b-30fc-4443-97e9-829742ba1e62 · outbound

This paper cites Loss aware post-training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Loss aware post-training quantization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.970339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.545529Z digest=sha256:2bd6b5c236701a75d0f64fbf64bb79bf5cc54d67cebb8573639d0d61d07edcc1

Observation 080915ef-394e-4e00-845d-b555b13a803e · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.953361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.551616Z digest=sha256:064554e442bba52703f5a2d1c0040f1773d740bfe488c39671e78dd7a0c03ebf

Observation 5e6e918e-5112-4acb-8998-657589850c16 · outbound

This paper cites Up or down? adaptive rounding for post- training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Up or down? adaptive rounding for post- training quantization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.935823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.556800Z digest=sha256:283d0805efa2359fc2d6af69128ac2c7d2f4242df482e35f8019bfdf928b9b57

Observation 26777c0e-02f8-4ef8-8aca-08f2bc01fa26 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.562492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.562492Z digest=sha256:5f75d42216015639a624980fabbdf80b177363dc12f02da32f2208de7eee52ef

Observation eb585fe7-2705-479a-84eb-424dbe2362c1 · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compression and acceleration,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Awq: Activation- aware weight quantization for on-device llm compression and acceleration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.918359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.569140Z digest=sha256:3d13b2831a1464e078aaeb00a3d31d1e4e5b3a9a516437309aecde0e215862f1

Observation d959deeb-80d8-4ca0-9b45-3fab8ec4f7a9 · outbound

This paper cites Q-vlm: Post-training quantization for large vision-language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Q-vlm: Post-training quantization for large vision-language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.895542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.575244Z digest=sha256:d45753f427b9a3eca604e0a4eb2462c22a1aa3939aacf68ce498c29d9dbb2b70

Observation d0470db3-9213-4550-b031-f30ff1dead43 · outbound

This paper cites Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.877051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.581813Z digest=sha256:0a5e5c181c5e36acbad60ff5f80f05197ed746a697629138ecbb67f4e5011111

Observation 71c0371c-cbd3-4427-9d06-8d0ae17c2a4b · outbound

This paper cites Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.857168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.588245Z digest=sha256:73c1baf04a8c36d940738bc28c38ae18251b5e89c38ec389b038d4a642437922

Observation 246f09ef-fe89-4379-8f6c-7cdabb71c9fb · outbound

This paper cites Ptq4sam: Post- training quantization for segment anything,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4sam: Post- training quantization for segment anything,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.834265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.594252Z digest=sha256:b07dfa59262ec79b6bde4a22d653e4df419e4ea01edafdbc6fe2d9dee480cea5

Observation 16077d71-deda-48f9-8cb1-53adac73f52d · outbound

This paper cites FP8 Formats for Deep Learning.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 Formats for Deep Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.600094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.600094Z digest=sha256:3aba0cc20dca9bad80a74cbd977da194397846313dc8ea23e6a4d9b45874da21

Observation 1bd83e60-759d-4ada-924b-038855fe6362 · outbound

This paper cites Faster Inference of LLMs using FP8 on the Intel Gaudi.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Faster Inference of LLMs using FP8 on the Intel Gaudi

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:06.326317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.606709Z digest=sha256:086a7e9e5ee3e8b482c49b7193db24ad96e03cff9dcad7a0497bbf6252fe67ab

Observation 838360a2-baa4-4155-99d9-7c9a44345f88 · outbound

This paper cites Fp8 quantization: The power of the expo- nent,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Fp8 quantization: The power of the expo- nent,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.814266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.613578Z digest=sha256:e342d0240759014ca71aae72b6ce080af44177d6eacb4df058c1588325fa5f89

Observation c798dbe1-7d59-4c39-bb16-ae3069122cd4 · outbound

This paper cites An Inquiry into Datacenter TCO for LLM Inference with FP8.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference An Inquiry into Datacenter TCO for LLM Inference with FP8

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.619868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.619868Z digest=sha256:6b95dfb157a52e44f9f236e123d7069d1b02d3eb8cf335a7d8f8a6f4aac19d3b

Observation 2eb5b9e2-2d15-4d7e-a109-12e46ba393d6 · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 versus INT8 for efficient deep learning inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.626348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.626348Z digest=sha256:3b07236910d85b63ed3331ff70e3b0533285a04ffd21520a917cfe2504f11232

Observation 737d520a-c7aa-464c-97e5-ec41adcffdce · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.631849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.631849Z digest=sha256:b17c73308ad085a51fa5c07f4002a48540a402369c40a6bb1b2e9f2e30fd972b

Observation 2041b6a0-c748-49ed-bdc0-c59c99c7781a · outbound

This paper cites The Super Weight in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Super Weight in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.638811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.638811Z digest=sha256:eee2b1ed99f55424271199014b5ba33a3e810d6026afa25bd8b63ed0f4b0ec32

Observation 256a2434-d6c1-47b5-af54-542cc5f7bff8 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quan- tization for large language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Smoothquant: Accurate and efficient post-training quan- tization for large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.792801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.645160Z digest=sha256:9a927b0f05b91726b8007a95fe99caafa4fb8081b0f7f3ec991aa38b43ea210b

Observation 7ff6de07-f271-49c8-89c7-15423dd1ddcc · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.650647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.650647Z digest=sha256:98a7988293c3055d29f9a51004a27f02fc9e9e9c0b34066817839b2174c2ec3b

Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.656359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.656359Z digest=sha256:f020c6e42decaff12f2b02ac6c7cc56a79b98a1dba631e73607087b84452d00d

Observation 3930d6bd-754d-4e32-b8df-b762c1043d1b · outbound

This paper cites Towards accurate post-training quantization for vi- sion transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Towards accurate post-training quantization for vi- sion transformer,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.774110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.661934Z digest=sha256:92159d0eaff60515077cc6c001a215c59436ff1032316e133de7d11b35d0d1af

Observation 95cf1fa8-c207-4e47-bc0a-7d556ba4f0a5 · outbound

This paper cites Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.752996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.666862Z digest=sha256:fc8779cfde6be4ae723c838f4f5e9fd6eea45e05f18e1804258489cefdaa6d66

Observation 6c4e8709-e7a7-4e5d-9709-47d47966f795 · outbound

This paper cites Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.672891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.672891Z digest=sha256:022632ca6ae77d7945616df7e9e892f90b0512b13c04fa18f0df958dba9c8ca5

Observation 0f4a9bee-0eb8-4885-ac3b-a90dbc2e4367 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.678322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.678322Z digest=sha256:333e40d24b07da5528769acbb559fec8340f536c911659cce999a131520b07a5

Observation 288a4b30-0a7e-44a2-9498-3fe66a8649ad · outbound

This paper cites Half-quadratic quantization of large machine learning models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Half-quadratic quantization of large machine learning models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.730280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.684564Z digest=sha256:a7a6e9c115b5cd55f95539173d6a25f0f1c0ba47b9ee2b343f68ad26228fea6d

Observation 0acc76a7-1570-4324-9422-4a8600415e8c · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.706490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.690691Z digest=sha256:f44cf25e1a90513ddb11c75493100a2136b58cd307e679483d38fa94bf1de542

Observation 93cd2ddc-5142-4fe7-b545-2e32eeb6e68b · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.686680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.696624Z digest=sha256:6c37911689118167e554e4c62cc0b5061e836d93d76176a1ca2f4d51f6e2798d

Observation 1bc2262d-b2c5-4f0c-ba8e-28df625beb99 · outbound

This paper cites Optimal brain sur- geon and general network pruning,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain sur- geon and general network pruning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.667965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.702137Z digest=sha256:9c2d5c7c3bcd87972df2b7ed8de07279db24bda281bd0007998c26f92d40f837

Observation fa35d4cf-af9e-4a6c-b26c-467877ef1d27 · outbound

This paper cites Optimal brain compression: A framework for accurate post-training quantization and prun- ing,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain compression: A framework for accurate post-training quantization and prun- ing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.647598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.707688Z digest=sha256:42f3434a2369a4bf1a34e4c4a0e88241c6a65a543ddf84d6410e0c8c527bef31

Observation 24f09a99-4a9d-4cae-bad6-bd6e87e8ce7c · outbound

This paper cites Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.623799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.715369Z digest=sha256:fc62e8d062119487d88fa0d029b9dffd0aa9b21564b3b4d70f9b27398fb82307

Observation f168e218-e86b-4e7b-932e-e33431ef6247 · outbound

This paper cites QQQ: Quality Quattuor-Bit Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.720980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.720980Z digest=sha256:2440f636c27ee9f71ea733f95a27f7d505a72190e1f0be9d13300eade027a2e5

Observation fd3c5138-7ee3-4af2-8310-3fa74fe23239 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qlora: Efficient finetuning of quantized llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.601626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.726742Z digest=sha256:7c72bda77b171e8af3d1f13fc2ec17082b7dd303b77f1da874dc1ae9d7bcef6e

Observation 45deaeb0-6ad2-445a-a43a-de4b52f0b1e6 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SpinQuant: LLM quantization with learned rotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.731877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.731877Z digest=sha256:7c12b555d2c405ed24e1ba10a6263adeb2b3dd27ccce31e4526baab477419833

Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · outbound

This paper cites Massive Activations in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.737732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.737732Z digest=sha256:cf69c18ddeefa8204a5700a6177e734e8357775f75d93a3def7d9dde8981cb11

Observation b8afbdc7-d703-48ec-8fb4-9b0bbc7b8aec · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.581618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.743472Z digest=sha256:e1902bfb3aa09ad44b297de98e4b644a80e8969c3f4aaabde1118943cd02ca02

Observation 4574bcbf-a4c9-4416-8eea-b5ec034b1866 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmbench: Is your multi-modal model an all-around player?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.561697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.748996Z digest=sha256:cec944d6c4dfa23765edededca2247bedaebfa2458a45c0f8fd86e2d70c4ab14

Observation 3745ee05-ec80-4ae7-a844-0a136a981a5c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.754358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.754358Z digest=sha256:e15dfb09449cfe47b0b588099320b5e1e3d59b8dc1b26ec2765690a08ff83d95

Observation 00b31de8-d9bc-4270-b39f-07fc5ad2812a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.760223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.760223Z digest=sha256:c9c6b399e10d438fb0c948df6b2398b69ce3a92ebeca576d45d38eaf63a35043

Observation 42cebc4f-9923-4baa-be4b-2fc40dc95cfb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.539573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.765099Z digest=sha256:c8cd7d580bea37548bbaab80c6d5fdc26026adc8c2a95b2cdae7e805a25f89f1

Observation 4bfbd2ce-a4ad-4242-abdb-26744edb88a7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.771908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.771908Z digest=sha256:2b0c2b2f0c17bd9a5b3a25b30fcad028d528027c0c2dc2ff4adf21dac82de4b5

Observation 7e06e537-d460-4808-8db1-67c7902ccab0 · outbound

This paper cites The Llama 3 Herd of Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.778165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.778165Z digest=sha256:9a85bdd57af54c7c3921446c06bc0892a256ec242fbde00d3666c0f209c6c7e1

Observation 6344a6fe-09f4-452f-8914-209142fd5ef1 · outbound

This paper cites Pointer Sentinel Mixture Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pointer Sentinel Mixture Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.783715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.783715Z digest=sha256:5cfa287bf241a8a9771118eac004a5108243d904556829466504a49343d9845e

Observation 7c1037a0-f90d-49be-a677-3cf71978fae8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.790299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.790299Z digest=sha256:62e9765c02e95ed552594ab25a5baa2c22e8580ec87286f9dc69b43c51ca66f6

Observation 2d12ea00-6193-4ffc-83d4-73c7f518e560 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.797383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.797383Z digest=sha256:412a11a4fefaec517842cdafbb366b69f778ab469d64662d3e28113b392577b4

Observation bff160fa-db2b-4996-9a0b-bc3ed1d1035b · outbound

This paper cites Language models are unsupervised mul- titask learners,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Language models are unsupervised mul- titask learners,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.517796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.803578Z digest=sha256:58460fc818d4aacfebee6937820f5e536d4295e73d11ebc0087edd20de9a3748

Observation 200b87c5-d284-4a8d-9322-28264c0cb5e2 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.809663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.809663Z digest=sha256:e008f66ecb0ca8478460df63aa53313041ab57e9a4b11502fc35719724e81bdd

Observation d74b1fd4-21b7-48a3-be62-12be5560fa0b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.815524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.815524Z digest=sha256:bb52739e30ec48b904b2c9162e49e20800eb23088b21c300a407606fb5782d30

Observation 773700e1-cc7f-484f-9068-9c7681b95810 · outbound

This paper cites Reason- ing about physical commonsense in natural language,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reason- ing about physical commonsense in natural language,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.494418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.821393Z digest=sha256:58a20e4663dac41a156861c151a5c5f521bf7a7cdfaab9774693b3023717e8d0

Observation 307dd09e-be6c-4d85-8b7f-2c63168013cf · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.826444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.826444Z digest=sha256:6cb6663775af8eaadb5fffac850fd5dde7015b5d6ca6fe1d94539cdecd7aed20

Observation d2267353-2366-4298-b1d5-bf7637c81f46 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.831954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.831954Z digest=sha256:f1dc4c86861b7d355c128e3851dc3a6747ef5adfa8cadafabe14d4aac22c6137

Observation 14025751-8455-4e36-a77e-90ac1c071fa5 · outbound

This paper cites A framework for few-shot language model evaluation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference A framework for few-shot language model evaluation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.476599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.838070Z digest=sha256:35b616e78ce09f98f9d042ad304cb02845e06b156d926e2b23679e1171154f11

Observation 3973de9a-8f95-417c-9b88-ddac976c38c2 · outbound

This paper cites Semantic parsing on Freebase from question-answer pairs,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Semantic parsing on Freebase from question-answer pairs,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.455038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:36:05.843955Z digest=sha256:502ad8c00af246a385e9e108cb4c234eba5a4c99debc66482b9fa19487e99d66

Observation 82a53ce8-7103-4aba-baf1-5aa3ba9503b1 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.850300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.850300Z digest=sha256:49210d6d646e7295678cef0f9ec1887988c17e84ff5e71820f447e64d1f609e6

Pith citing papers

No inbound Pith citation observations are available.