Pith. sign in

Paper Citation Record · LEDGER

MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2408.11743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11743 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T01:15:20.402556Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:45.492485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 317a89ed-31ea-490f-b39a-6968f2e8019b · inbound

A Proximal Operator for Inducing 2:4-Sparsity cites this paper.

A Proximal Operator for Inducing 2:4-Sparsity MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T01:15:20.402556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T01:15:20.402556Z digest=sha256:b9e9e1ba6cd7cf052a72ef8b55d3ca52c62c01a75c79688d948030ab9b44dd6c

Observation 3d38cf8b-ea40-411f-9334-1878b369d20b · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:01.423333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:01.423333Z digest=sha256:f1194225688add23e45e806f5e4bdc32457e610c03052151deac33a0acb38c1e

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:fdc41d4d3d507e2b2679d18e77122df5305cb8c3465c59a729356f88d30de32c

Observation fe734b39-930f-4176-a525-f89adae7241c · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:32.827124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:32.827124Z digest=sha256:dd7a48eeb05a4e775a4ee7af8efe83b6d5c15baa9dcae46ff721e3e654e4f8de

Observation dba3f5c5-a497-44c2-aed6-0a5112f34dd6 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.035971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:7548d1c607b2454bdae280de5306ca653e69f65f4d7800068f8c002460aa13e5

Observation cc2268d3-fe29-4278-bbfc-c56238e32119 · inbound

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate cites this paper.

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.871774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:15:25.783283Z digest=sha256:23d142a1df28cb9de3abb652bea61d0898ff21c6ef068c82ff87841f040609f2

Observation 748daa29-4c86-42e9-b226-daa366409178 · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.037549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:e90e1010abad6f46d31987731688df3204360c32662ff18594f1c86eb3109d16

Observation 0383d94a-7265-4073-a26f-8b039f84430d · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.389954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:457e0b7f66f47cd0bdd9055d58bb44ad12c9fd1e94671aa40a53dec623933f2e

Observation 0b29538f-c8c9-48b2-a311-c077cbef8d77 · inbound

Statistically-Lossless Quantization of Large Language Models cites this paper.

Statistically-Lossless Quantization of Large Language Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.085341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T16:22:13.399819Z digest=sha256:489a76550c128a89ced3a45de28f64e51522b4e6d422a87fdcefb48d86f8ddb9

Observation 00ce7d43-f1d6-40fc-80e7-5d319d9656f2 · inbound

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference cites this paper.

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:58:34.202730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:57:18.944617Z digest=sha256:17f4970321a958b042bb793b3b0718dc7dd09ba28ec3147464a4207ad1d7eaf3

Observation 99b5642a-a6df-4c1d-8d29-951f569a962d · inbound

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models cites this paper.

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.737221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:41:57.350416Z digest=sha256:cf18502c2a1a48e80c3573651405c4ed41ba70b51a7e865b5aa092be1f6a01d4

Observation dff7ba71-ad0f-482c-8229-0f5facc5dc09 · inbound

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis cites this paper.

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.436235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:59:11.821440Z digest=sha256:0fa10dc080f3a80d348f99b6c953b128b21e1185373435c52eb4c0257dc32d69

Observation 0387748f-437f-4a66-821c-29a58aafe98a · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:57:44.946716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T11:26:56.230283Z digest=sha256:d1972234eb109d5f8afc4a61b96f011bd8e16396ca6b7d906c8f71ef6ea42f58

Observation 3a4a5887-8b73-41cd-8e0f-03d3901bdf78 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:02.894834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T23:37:25.391853Z digest=sha256:ece1b5fcc5edd5e43c1a2b9f760e1b7557e808fd1bbc60571270c779e6328639

Observation 41709bc4-e638-4950-a218-3121d2e44dbc · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.493866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:379ae3abfe2908d091033b91fbc74970d15ac7e3bc098fdedd875a2a8edb57c6

Observation b43e328a-fb3f-472c-96ff-8c1c3a949f1f · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.293856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.293856Z digest=sha256:beae64292b0f3ee4b4872a555926f41abe4195e5ab8134e5aac145e8545ed933