Pith. sign in

Paper Citation Record · LEDGER

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.14638.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14638 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.850300Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 236c17b8-9918-47c9-aa7d-3e09aa1d7f80 · outbound

This paper cites Learned Step Size Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Learned Step Size Quantization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.491210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.491210Z digest=sha256:9132b217fa77473a132b3ce8c0f7e1397f405c67f8fc7d8c7a150b51b171eb4a

Observation 94b1badb-5fc8-4880-ac05-496904854c23 · outbound

This paper cites Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.082662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.498781Z digest=sha256:653cae6783c5d51d545fffabed806594d927d8a15ac6a67e03fc3a9dd86feeca

Observation 2c1403f1-ef5d-45e0-9004-6e41e8b4630e · outbound

This paper cites Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.056103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.506094Z digest=sha256:097fddce6f4423bd13935bbe306cbe6edfe5ef9ec74ca5398dae00c7fd5329ce

Observation 1c5fd896-d33d-4a13-8799-4efc3d93369b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.513839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.513839Z digest=sha256:99b40c77bc168d19e0127c5adc82944771e28f0f860c5e9027dac8d2ab760795

Observation 89a56030-cdbd-4dc1-94f1-e2a0167bf1b3 · outbound

This paper cites Data-free quantization through weight equalization and bias correction,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Data-free quantization through weight equalization and bias correction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.030194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.520642Z digest=sha256:d91ed82b8cb6b29be9ccb7964c9e101afa9067784f2ca632ead1f945c195402e

Observation e0c27305-2a9d-404e-a070-7be9039062a8 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.527071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.527071Z digest=sha256:a4f182716b322be0a7d43ef43fe33a57060392151dca33293a40bec2f45c23ff

Observation 13814e29-fb9e-4965-9a5a-ca9090c6622c · outbound

This paper cites Accurate post training quantization with small calibra- tion sets,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Accurate post training quantization with small calibra- tion sets,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:07.009418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.534639Z digest=sha256:f7cf69bb70a9da7347dd1d2fec2e44d191df65611464d7c64245f954b7aa877a

Observation c581bb25-a830-4af4-a641-a331f5e4e420 · outbound

This paper cites Post-training quantization for vision transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Post-training quantization for vision transformer,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.987699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.540267Z digest=sha256:236c0f0a78ae7ffd0c0e114e4670667a638968c597c6319b3764dde326138556

Observation 0475f67b-30fc-4443-97e9-829742ba1e62 · outbound

This paper cites Loss aware post-training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Loss aware post-training quantization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.970339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.545529Z digest=sha256:c41cc21ed3473cd0d50b157273e2a3eafd9892165d6d8011f127e371c5cae88e

Observation 080915ef-394e-4e00-845d-b555b13a803e · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.953361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.551616Z digest=sha256:b2435f5469e8e50dda1451bc1dedb09f1fe29194b9a5dcd4578350a5a75abb01

Observation 5e6e918e-5112-4acb-8998-657589850c16 · outbound

This paper cites Up or down? adaptive rounding for post- training quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Up or down? adaptive rounding for post- training quantization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.935823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.556800Z digest=sha256:db8c8de14db14dae2f4a7fc0c005ef16587843a47fea04ced527526d6b476690

Observation 26777c0e-02f8-4ef8-8aca-08f2bc01fa26 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.562492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.562492Z digest=sha256:68708695cec14f10fa344471e006d327012bcb472ffdd3f4562b1d1b48915ae5

Observation eb585fe7-2705-479a-84eb-424dbe2362c1 · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compression and acceleration,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Awq: Activation- aware weight quantization for on-device llm compression and acceleration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.918359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.569140Z digest=sha256:0a150a4cb2c96b392815bb2f408ceb37307db4f1bf87492d38548c642c17267b

Observation d959deeb-80d8-4ca0-9b45-3fab8ec4f7a9 · outbound

This paper cites Q-vlm: Post-training quantization for large vision-language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Q-vlm: Post-training quantization for large vision-language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.895542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.575244Z digest=sha256:f280524987464323ff6000b8b2bc2617b06c8385a73642aeb25a7a96e811c5e6

Observation d0470db3-9213-4550-b031-f30ff1dead43 · outbound

This paper cites Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.877051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.581813Z digest=sha256:efe36c033cc9498f83ecf5e6aeea483c761d5137942c34038cbb78576b3eb9a6

Observation 71c0371c-cbd3-4427-9d06-8d0ae17c2a4b · outbound

This paper cites Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.857168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.588245Z digest=sha256:75ff3a87d083098e2a3a8214c0e6452434b5ea8bb6847275c2a10f0a6618255e

Observation 246f09ef-fe89-4379-8f6c-7cdabb71c9fb · outbound

This paper cites Ptq4sam: Post- training quantization for segment anything,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4sam: Post- training quantization for segment anything,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.834265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.594252Z digest=sha256:8eb70b2edb591f2d5483a1dfc6cffa3b438c13e0e97e2971341a0f33491a95a5

Observation 16077d71-deda-48f9-8cb1-53adac73f52d · outbound

This paper cites FP8 Formats for Deep Learning.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 Formats for Deep Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.600094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.600094Z digest=sha256:e85cfe28a6a6c4d6fa3c2f94034f8712513ff711a3d9a883d0322040b2cf2342

Observation 1bd83e60-759d-4ada-924b-038855fe6362 · outbound

This paper cites Faster Inference of LLMs using FP8 on the Intel Gaudi.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Faster Inference of LLMs using FP8 on the Intel Gaudi

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:06.326317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.606709Z digest=sha256:2254738325ee2db653ebda81436c1d04aa4420c98835bed09fd9bb3ddd4c12d7

Observation 838360a2-baa4-4155-99d9-7c9a44345f88 · outbound

This paper cites Fp8 quantization: The power of the expo- nent,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Fp8 quantization: The power of the expo- nent,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.814266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.613578Z digest=sha256:3b0d646972d8b86c79f959b072683a9322b4f9e71fe1cd8a6fb715b160cd7c60

Observation c798dbe1-7d59-4c39-bb16-ae3069122cd4 · outbound

This paper cites An Inquiry into Datacenter TCO for LLM Inference with FP8.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference An Inquiry into Datacenter TCO for LLM Inference with FP8

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.619868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.619868Z digest=sha256:85389adca2c2c437d302fd7902c003c6ce2e21fe754d1e2b23ec5e1f84db3ad0

Observation 2eb5b9e2-2d15-4d7e-a109-12e46ba393d6 · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 versus INT8 for efficient deep learning inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.626348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.626348Z digest=sha256:275cbd8800c5bafaac18bb6fef79bde05d01ca5f18134c666fb27d0210d0b04f

Observation 737d520a-c7aa-464c-97e5-ec41adcffdce · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.631849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.631849Z digest=sha256:71a62e00cf3dc88fe0aa320304084c670295dc30ecf29bc664804b889991d871

Observation 2041b6a0-c748-49ed-bdc0-c59c99c7781a · outbound

This paper cites The Super Weight in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Super Weight in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.638811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.638811Z digest=sha256:40097047e4ef3fb64e7a18190eaa45d72a73db79e93aa47c9bb2b2a0c5edd169

Observation 256a2434-d6c1-47b5-af54-542cc5f7bff8 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quan- tization for large language models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Smoothquant: Accurate and efficient post-training quan- tization for large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.792801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.645160Z digest=sha256:5d4795c5ddb90fea464d6040607153e019e5b4db4b63d28013c7d393a124c40f

Observation 7ff6de07-f271-49c8-89c7-15423dd1ddcc · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.650647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.650647Z digest=sha256:77e031e210fa3de350e4425f6ffe006c50e038b964d3d56219d22146d43d4a3e

Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.656359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.656359Z digest=sha256:4ec7657a63b22247dc9a464b8338cf43ed335bc9789bd59188df517a8284f1c9

Observation 3930d6bd-754d-4e32-b8df-b762c1043d1b · outbound

This paper cites Towards accurate post-training quantization for vi- sion transformer,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Towards accurate post-training quantization for vi- sion transformer,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.774110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.661934Z digest=sha256:32edce0ca8cd5338b8533dcdce075f15e5e1d3e4c1555f109744e5ebc05434d4

Observation 95cf1fa8-c207-4e47-bc0a-7d556ba4f0a5 · outbound

This paper cites Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.752996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.666862Z digest=sha256:e50d85e08684993839a899b4cbd8a9455e884a221e19f55d31069c879e776d6e

Observation 6c4e8709-e7a7-4e5d-9709-47d47966f795 · outbound

This paper cites Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.672891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.672891Z digest=sha256:b83eabeaa5a00cd4f44bd14d1afff846cfab12f7198a14225b6a38ccd947775e

Observation 0f4a9bee-0eb8-4885-ac3b-a90dbc2e4367 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.678322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.678322Z digest=sha256:19cae30cb384d4c6d59ad85cddd90d161762aa91f4863800059a8273ea2432bf

Observation 288a4b30-0a7e-44a2-9498-3fe66a8649ad · outbound

This paper cites Half-quadratic quantization of large machine learning models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Half-quadratic quantization of large machine learning models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.730280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.684564Z digest=sha256:50403602916b5f85d91dded03e56436a5778de8a766fd8f517c31f968828f7f3

Observation 0acc76a7-1570-4324-9422-4a8600415e8c · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.706490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.690691Z digest=sha256:aeccd14a3410639fad54429dd710c60ad72d00c8d3f1cfb7b77d5c99f6f407a5

Observation 93cd2ddc-5142-4fe7-b545-2e32eeb6e68b · outbound

This paper cites Pd-quant: Post-training quantization based on prediction difference metric,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.686680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.696624Z digest=sha256:17f4b7502605dcaaa52e92472b6433ee603e032ba27d95e16f6872b2c558fd66

Observation 1bc2262d-b2c5-4f0c-ba8e-28df625beb99 · outbound

This paper cites Optimal brain sur- geon and general network pruning,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain sur- geon and general network pruning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.667965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.702137Z digest=sha256:d3da787ec24def30167ac453ecd18ce4ccec794d635fb727e831efee963cf253

Observation fa35d4cf-af9e-4a6c-b26c-467877ef1d27 · outbound

This paper cites Optimal brain compression: A framework for accurate post-training quantization and prun- ing,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain compression: A framework for accurate post-training quantization and prun- ing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.647598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.707688Z digest=sha256:a359e9dabf09e9fa99fdf5e555613dbe0b2d5033e6ea2b6fcea155490ac911b3

Observation 24f09a99-4a9d-4cae-bad6-bd6e87e8ce7c · outbound

This paper cites Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.623799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.715369Z digest=sha256:52a060c149d98e91385d80929a29e79d32987707985cced63b205c16dc4c2e3a

Observation f168e218-e86b-4e7b-932e-e33431ef6247 · outbound

This paper cites QQQ: Quality Quattuor-Bit Quantization for Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.720980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.720980Z digest=sha256:648f6b364d4b4651ecaa53ec932df379a2b6b4f6759eda932f9b8bb7a3a17125

Observation fd3c5138-7ee3-4af2-8310-3fa74fe23239 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qlora: Efficient finetuning of quantized llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.601626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.726742Z digest=sha256:84cecc0b81aec1c3d48480bf1014214c6161f250f153a58ed62e98b133237621

Observation 45deaeb0-6ad2-445a-a43a-de4b52f0b1e6 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SpinQuant: LLM quantization with learned rotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.731877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.731877Z digest=sha256:3955319044dbd0516ab0b3c8280db8c2db81698d2dfbb6ad828e42123d9a8e31

Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · outbound

This paper cites Massive Activations in Large Language Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.737732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.737732Z digest=sha256:85ebabc780a7e8dd091f232234ba44bfcfb1d31fae8eb17eb9917c1ec76b2f42

Observation b8afbdc7-d703-48ec-8fb4-9b0bbc7b8aec · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.581618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.743472Z digest=sha256:321b96b0c78180285103b7a74c18fc9563d2e39387fe10d9b8f1c948906c132e

Observation 4574bcbf-a4c9-4416-8eea-b5ec034b1866 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmbench: Is your multi-modal model an all-around player?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.561697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.748996Z digest=sha256:5507440052315ae1e53649febf200ace35fb44e9aff21486e8f72c9c019d4801

Observation 3745ee05-ec80-4ae7-a844-0a136a981a5c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.754358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.754358Z digest=sha256:e9ae5cc05f17306b7a986fb6419b86258b1de644cec96fcdf0b8f8eb4643c360

Observation 00b31de8-d9bc-4270-b39f-07fc5ad2812a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.760223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.760223Z digest=sha256:1147c108430b63272a2de204109e6f927f4b0292dc3bc26acf304c7906eae510

Observation 42cebc4f-9923-4baa-be4b-2fc40dc95cfb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.539573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.765099Z digest=sha256:5dbac9a4be24e4319e67286c88fb73c31a83e571f24db101b317a659b82fdb29

Observation 4bfbd2ce-a4ad-4242-abdb-26744edb88a7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.771908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.771908Z digest=sha256:77e807f563cd182d19f71207ce356713d77f7be73f903c45932e8e667821251d

Observation 7e06e537-d460-4808-8db1-67c7902ccab0 · outbound

This paper cites The Llama 3 Herd of Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.778165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.778165Z digest=sha256:787f674fa8700d08d4503479fa3113dc0bd1e192b0dbd9f209b796ce175f9d88

Observation 6344a6fe-09f4-452f-8914-209142fd5ef1 · outbound

This paper cites Pointer Sentinel Mixture Models.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pointer Sentinel Mixture Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.783715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.783715Z digest=sha256:6dec4475ca04bf072fbc296ad2046785552bd2ba0b6b57035baf4cec85fa22cd

Observation 7c1037a0-f90d-49be-a677-3cf71978fae8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.790299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.790299Z digest=sha256:8542acf608be4c34a1da3be1474a145aa59c6b7749c84600847106034d90a6eb

Observation 2d12ea00-6193-4ffc-83d4-73c7f518e560 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.797383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.797383Z digest=sha256:627bec920347cb3942ec8cbc07f10b9f80e45b0135e70666ae3ce6e1e0755a30

Observation bff160fa-db2b-4996-9a0b-bc3ed1d1035b · outbound

This paper cites Language models are unsupervised mul- titask learners,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Language models are unsupervised mul- titask learners,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.517796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.803578Z digest=sha256:c30f7339ed0f39f2c476e7ffbf139dcbbd232b32d22dba65cf92063167090fff

Observation 200b87c5-d284-4a8d-9322-28264c0cb5e2 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.809663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.809663Z digest=sha256:276c10425ab22dd3c7411725ebf7a3f34a187cff5bf9d263fecc5fb2af84be17

Observation d74b1fd4-21b7-48a3-be62-12be5560fa0b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.815524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.815524Z digest=sha256:3b8df9c69c2eabddf1c92e99a5c8b09fea0b4fe52810804a897f4705bfd06bc4

Observation 773700e1-cc7f-484f-9068-9c7681b95810 · outbound

This paper cites Reason- ing about physical commonsense in natural language,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reason- ing about physical commonsense in natural language,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.494418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.821393Z digest=sha256:cb4c17c9c35078c0ba9f6ae40b6d868468df582add28cbf9ce22d7559a0817f8

Observation 307dd09e-be6c-4d85-8b7f-2c63168013cf · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.826444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.826444Z digest=sha256:92c49b8f17a61c23ce71ef23fb385148b76410c4bb18bf0079c78b7faa52d2c7

Observation d2267353-2366-4298-b1d5-bf7637c81f46 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.831954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.831954Z digest=sha256:a6931263d87b056362a6a2bc35034951a199e95c7959abdf65594aefda0fed7a

Observation 14025751-8455-4e36-a77e-90ac1c071fa5 · outbound

This paper cites A framework for few-shot language model evaluation,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference A framework for few-shot language model evaluation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.476599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.838070Z digest=sha256:dcff42cbbc91b1957bcb9e31217634522bc03d766ba5ef43fb4007ca31f50e16

Observation 3973de9a-8f95-417c-9b88-ddac976c38c2 · outbound

This paper cites Semantic parsing on Freebase from question-answer pairs,.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Semantic parsing on Freebase from question-answer pairs,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.455038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.843955Z digest=sha256:568c8b631d615dd873099b707b12c3a21918a3db0f91d9cf978c8a90ab5a0153

Observation 82a53ce8-7103-4aba-baf1-5aa3ba9503b1 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.850300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.850300Z digest=sha256:8a96fda6d2df23eaf29d8abdf62e8cca5855afa5098388944fff133012888a43

Pith citing papers

No inbound Pith citation observations are available.