Pith. sign in

Paper Citation Record · LEDGER

MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2408.11743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11743 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:14:01.423333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:45.492485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3d38cf8b-ea40-411f-9334-1878b369d20b · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:01.423333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:01.423333Z digest=sha256:fd0f7eee3bd6140c636ba829f240bbd0844e2acf88536c2f6b999019d8178870

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:fdc41d4d3d507e2b2679d18e77122df5305cb8c3465c59a729356f88d30de32c

Observation fe734b39-930f-4176-a525-f89adae7241c · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:32.827124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:32.827124Z digest=sha256:dfd567a169f574744beadb40a9cedb8ea602b233c094972f2111d85a00ac7894

Observation dba3f5c5-a497-44c2-aed6-0a5112f34dd6 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.035971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:1bfc5c3d1b4a965088c6372002168a25622ded457dacc85e38ba33ca607f3fbb

Observation cc2268d3-fe29-4278-bbfc-c56238e32119 · inbound

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate cites this paper.

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.871774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:15:25.783283Z digest=sha256:df0cdd7326e1618e7a6a0f96ea16eb30a0f12352dab9a7a97a4d2d6082e4240c

Observation 748daa29-4c86-42e9-b226-daa366409178 · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.037549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:5ecb55400313f61d69440a2a6b6256d7a03e08fef0682dd91d745b9ae19b41e0

Observation 0383d94a-7265-4073-a26f-8b039f84430d · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.389954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:21ac0ebc2a5dc40dd324e8f576844da40e345fc3ce55d4f11a0d2e49709d5cd5

Observation 0b29538f-c8c9-48b2-a311-c077cbef8d77 · inbound

Statistically-Lossless Quantization of Large Language Models cites this paper.

Statistically-Lossless Quantization of Large Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.085341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T16:22:13.399819Z digest=sha256:872a20510da9e80fd1c81057cb38dd1edc24c8d5b6d625964d16e52a580e8a85

Observation 00ce7d43-f1d6-40fc-80e7-5d319d9656f2 · inbound

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference cites this paper.

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:58:34.202730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:18.944617Z digest=sha256:e3ad434c55a7c993161fdb44bbd1baf495e41a0bf3d1bf93eb54a4d8ab506c33

Observation 99b5642a-a6df-4c1d-8d29-951f569a962d · inbound

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models cites this paper.

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.737221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:41:57.350416Z digest=sha256:d750731b544aafb746174be367d1ae84201987d2748877e55bb527abffafade7

Observation dff7ba71-ad0f-482c-8229-0f5facc5dc09 · inbound

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis cites this paper.

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.436235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:59:11.821440Z digest=sha256:34939436c565f1194d0267bf6042e57c65446c37d25490002b1170ea5328c946

Observation 0387748f-437f-4a66-821c-29a58aafe98a · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:57:44.946716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T11:26:56.230283Z digest=sha256:c08b33a4fa3de036627510ed9d42d597b537c16ee9a501fccf9c373d30cbd155

Observation 3a4a5887-8b73-41cd-8e0f-03d3901bdf78 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:02.894834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T23:37:25.391853Z digest=sha256:1ba8bda5da10c53a8bb89056a7cf8638f886266d272be812ea66c8be834aba12

Observation 41709bc4-e638-4950-a218-3121d2e44dbc · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.493866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:4b71faa7e1efbbb7b203eac4f7af77a45c4bd2c40b5390f49ebbcbd0f777015e

Observation b43e328a-fb3f-472c-96ff-8c1c3a949f1f · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.293856Z digest=sha256:beae64292b0f3ee4b4872a555926f41abe4195e5ab8134e5aac145e8545ed933