Pith. sign in

Paper Citation Record · LEDGER

Retrofitting Linear Attention into Diffusion Language Models

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.06628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06628 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:14:52.259504Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a2cd5a-19a3-4a67-9705-f5aaf2f93008 · outbound

This paper cites LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.215311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.215311Z digest=sha256:904ad378e91ed917ab468b8b5f1d89617c984e95a42ae7c0e7b4e5d62a48e9f1

Observation a3bd89b5-400f-4c24-ae2a-c637e226a8f3 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Retrofitting Linear Attention into Diffusion Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.227519Z digest=sha256:cb41343b4a7b54ae0efb14eaa6aba16b4d8d6aab2a7e35f2c1a0e15ca669f6a3

Observation f2c19505-7d58-4ab0-8f26-18f9c71be134 · outbound

This paper cites The diffusion duality.arXiv preprint arXiv:2506.10892,.

Retrofitting Linear Attention into Diffusion Language Models The diffusion duality.arXiv preprint arXiv:2506.10892,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.234749Z digest=sha256:bb40879393b5c3629e7645089d4f7884d27a02a0b297a6dc5481b3f8a225765e

Observation 23b64969-f29e-4ae8-bee2-6ee729f11043 · outbound

This paper cites Simple guidance mechanisms for discrete diffusion models.

Retrofitting Linear Attention into Diffusion Language Models Simple guidance mechanisms for discrete diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:14:53.431302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.237955Z digest=sha256:1952818c72803545fea851227591096e1909fb2579a812f163c33843dfdba139

Observation 45fde0ac-3bb3-45ff-9b5a-00564e33aa02 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

Retrofitting Linear Attention into Diffusion Language Models Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.241127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.241127Z digest=sha256:7f94b37a95389b3dbe481d89a1b249badf8815cf9654cf0b8a56d0913b3db534

Observation 75a91339-753f-4b2c-ade5-355164e84f6f · outbound

This paper cites Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,.

Retrofitting Linear Attention into Diffusion Language Models Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:52.860526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.244505Z digest=sha256:18a9b71f74bef89633d45dded3a43ba81ee4e3ce2b8bdbcc752a959a9016b441

Observation b661e5b2-1b26-4a38-8fd4-2207cde2dd89 · outbound

This paper cites Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning.

Retrofitting Linear Attention into Diffusion Language Models Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.251220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.251220Z digest=sha256:324fe5d8add2e3501863537f036e7a2c804658791f4521c9e88ca9e29453f5cb

Observation dd92681b-0e94-4a2c-b337-b9e4a095f30a · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Retrofitting Linear Attention into Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.255473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.255473Z digest=sha256:bd90d0a9698f87d9e0532362a30f5f1e72ebcdaa549c6cf1cfd11d56d8fd25a4

Observation f6e50143-e983-4ddc-aab8-923000149b22 · outbound

This paper cites LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.259504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.259504Z digest=sha256:249d2445118b41129ffbb177c53c49e135abdb5906fa7246b0af1191fd7a4c2c

Observation b4fc0e3f-e488-4920-94dc-c5359c886ca1 · outbound

This paper cites Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions.

Retrofitting Linear Attention into Diffusion Language Models Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.223396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.223396Z digest=sha256:017fbddeeb13727b6be84042ebce3cfd6b40695b0ab2f239b2d23d3d417339e4

Observation 57eac887-b242-4466-82b9-7facf6251a04 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Retrofitting Linear Attention into Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.219401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.219401Z digest=sha256:122ee503253be3957c95398a84c6349091b2e5de29f60186ce80c55f75dcd467

Observation 4116384a-676a-4adb-beea-c56fef673046 · outbound

This paper cites an unresolved cited work.

Retrofitting Linear Attention into Diffusion Language Models Unresolved cited work

Reference 2024

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:53.207417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.231187Z digest=sha256:9c979818ee1d16a0aa4afdefe6b1fa07433e54cf6b231a70d5616ad7fb4fdfa5

Observation 005303c5-fdda-42e1-911d-718910c9471c · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.210581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.210581Z digest=sha256:f73741f0be81dd7f2efeac921c58b072dfa297c2c276502344da474015f03b16

Observation 0b87d0c0-a694-4142-bb9f-d2dfbc6d6f56 · outbound

This paper cites Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,.

Retrofitting Linear Attention into Diffusion Language Models Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.247619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.247619Z digest=sha256:a572c76872478a5340f0a42b9aefe4a450ef08d00517fcb730675f8a4262a2a8

Pith citing papers

No inbound Pith citation observations are available.