Pith. sign in

Paper Citation Record · LEDGER

Retrofitting Linear Attention into Diffusion Language Models

As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.06628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06628 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:14:52.259504Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a2cd5a-19a3-4a67-9705-f5aaf2f93008 · outbound

This paper cites LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.215311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.215311Z digest=sha256:62710a0ef5f45fa6bcf1234a841cab4fd193e69ca6e070351e8d866c430f8072

Observation a3bd89b5-400f-4c24-ae2a-c637e226a8f3 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Retrofitting Linear Attention into Diffusion Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.227519Z digest=sha256:4b8db1ea0916ab25009dfd0066ee7a8e9e11c91bcf026cf4471cdec1d4d31ca0

Observation f2c19505-7d58-4ab0-8f26-18f9c71be134 · outbound

This paper cites The diffusion duality.arXiv preprint arXiv:2506.10892,.

Retrofitting Linear Attention into Diffusion Language Models The diffusion duality.arXiv preprint arXiv:2506.10892,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.234749Z digest=sha256:de08e7b79d3041bea50a5d90e67eace818756335b37094e892e90abdb533b2c0

Observation 23b64969-f29e-4ae8-bee2-6ee729f11043 · outbound

This paper cites Simple guidance mechanisms for discrete diffusion models.

Retrofitting Linear Attention into Diffusion Language Models Simple guidance mechanisms for discrete diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:14:53.431302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T04:14:52.237955Z digest=sha256:110e6ac93a3ce8bddd252df0782a823bda3449f13ad8bd555d1332119f56a138

Observation 45fde0ac-3bb3-45ff-9b5a-00564e33aa02 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

Retrofitting Linear Attention into Diffusion Language Models Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.241127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.241127Z digest=sha256:9e268ef1fc3276c7b03f7c2323fee1fdc845568a7ac333193d65509a7314e614

Observation 75a91339-753f-4b2c-ade5-355164e84f6f · outbound

This paper cites Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,.

Retrofitting Linear Attention into Diffusion Language Models Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:52.860526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T04:14:52.244505Z digest=sha256:7ad3ffc3a6c338a33754809841e2082ca84f2809bed558d7d17f86f6e2f2b737

Observation b661e5b2-1b26-4a38-8fd4-2207cde2dd89 · outbound

This paper cites Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning.

Retrofitting Linear Attention into Diffusion Language Models Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.251220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.251220Z digest=sha256:48fd1647d1723b698b69d3cf84b8a8a619f562563abffb1c95316ee75c47d472

Observation dd92681b-0e94-4a2c-b337-b9e4a095f30a · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Retrofitting Linear Attention into Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.255473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.255473Z digest=sha256:dbc0d396264aea8216056c209da620deab093ea88df23a668a80af550003bdd8

Observation f6e50143-e983-4ddc-aab8-923000149b22 · outbound

This paper cites LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.259504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.259504Z digest=sha256:b55820e939199a2c927386c81c2945c43f7937655bdc394dd947dcdf12ea3d6b

Observation b4fc0e3f-e488-4920-94dc-c5359c886ca1 · outbound

This paper cites Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions.

Retrofitting Linear Attention into Diffusion Language Models Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.223396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.223396Z digest=sha256:686846cb181ded4abad31b390797cab13fcea7416e0bb18829ce0b943bd6a28c

Observation 57eac887-b242-4466-82b9-7facf6251a04 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Retrofitting Linear Attention into Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.219401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.219401Z digest=sha256:e40640f0a422e183468ddb3701c93598ba032251baf31130833c355ef0453a80

Observation 4116384a-676a-4adb-beea-c56fef673046 · outbound

This paper cites an unresolved cited work.

Retrofitting Linear Attention into Diffusion Language Models Unresolved cited work

Reference 2024

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:53.207417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T04:14:52.231187Z digest=sha256:35161fdd23d189e69d067c863800783378927718c736cde6463b5d7aae652fca

Observation 005303c5-fdda-42e1-911d-718910c9471c · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.210581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.210581Z digest=sha256:47326535e11aa629e73d97de44fb98dd9a9bab744102370b4de9f5a60990866e

Observation 0b87d0c0-a694-4142-bb9f-d2dfbc6d6f56 · outbound

This paper cites Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,.

Retrofitting Linear Attention into Diffusion Language Models Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.247619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.247619Z digest=sha256:db6704eeed254394913c125beb76c80de15779f9943260fed5369b6ff8d19f58

Pith citing papers

No inbound Pith citation observations are available.