Pith. sign in

Paper Citation Record · LEDGER

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.03340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03340 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:20:45.911627Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 388c28c0-eb93-44e9-99a8-8cdb77e13fa6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.849068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.849068Z digest=sha256:3dbb8926d3c40058cf444011a2a16b878db58581ba788912c6247c47cab132e5

Observation f861ce85-a9de-4744-a2b0-b7debd850260 · outbound

This paper cites Kasai, H.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Kasai, H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.670584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:20:44.987614Z digest=sha256:f05a108e3bb7da8ddbc05bc291e5d89f03d40e944f69c8b3db5faba3c14fe19d

Observation 78651683-8be8-495c-9bc2-70f2f9278298 · outbound

This paper cites direct” loss. This is natural because the cross entropy loss for next-token prediction is used in “direct.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency direct” loss. This is natural because the cross entropy loss for next-token prediction is used in “direct

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.259058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:20:45.911627Z digest=sha256:1ce98b5a0c041d51ee7491eb7ef58f1e0583ff88f2752b35d994d8d69e0236ee

Observation 9bcdbe42-d596-49eb-b26f-969247eb8437 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Linformer: Self-Attention with Linear Complexity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.766234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.766234Z digest=sha256:5895a34487603f2a357dec2431b8587c304799871aa45640f495156c02aab791

Observation 643d0de8-beb3-415d-bac1-0a91e45281a8 · outbound

This paper cites • K : Rd × Rd → R is the positive definite kernel given byK(x, y) = Ez∼τ [ϕ(x; z)ϕ(y; z)], where τ is a probability measure on a measurable setZ, and ϕ : Rd × Z →R is a feature map.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency • K : Rd × Rd → R is the positive definite kernel given byK(x, y) = Ez∼τ [ϕ(x; z)ϕ(y; z)], where τ is a probability measure on a measurable setZ, and ϕ : Rd × Z →R is a feature map

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.404280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:20:45.845845Z digest=sha256:2bef026bc74802f1fcb4cc0142871ad018ad21ce5c305d614b7c68f86cddff24

Observation 989f48a6-8f41-469b-9540-2095b8d385eb · outbound

This paper cites Scavenging Hyena: Distilling Transformers into Long Convolution Models.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Scavenging Hyena: Distilling Transformers into Long Convolution Models

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.341397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.341397Z digest=sha256:8bdc1d5e5f15be46b7ad96467c31e6adb7218e43d19476b0025d1abd64e70fca

Observation 96359c09-3a7e-4c60-820f-940e4e64484c · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.152141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.152141Z digest=sha256:9ed7a4adf47569cb35456ce3ce7aa658954f0452f9d6ced4ef79c5ff7ab6cb24

Observation e9165664-45c3-4da1-9cbd-34d955ad0e27 · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.689403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.689403Z digest=sha256:9c801976fcd0badcef385352707d1aae43a7b1551dee5ebb13e8ed4967143206

Observation 7cd2de6b-795d-4078-a062-e6e4f1d3cbac · outbound

This paper cites an unresolved cited work.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:20:46.799700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:20:44.742887Z digest=sha256:2ec4e2a03d2f09bffb9da77165837a0f5f98ec4890c8f19cff2db144fe37455e

Observation fba1e96d-fd8e-4112-97a3-ba5626d73da2 · outbound

This paper cites Sakamoto and K.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Sakamoto and K

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:45.609916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:45.609916Z digest=sha256:f5dce60a060967fdd51a9286f6344a11bb1f41caf4b8595a73586d5af743ba87

Observation 4920c157-511e-4a80-ad76-b027ccf4f504 · outbound

This paper cites DiJiang: Efficient Large Language Models through Compact Kernelization.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency DiJiang: Efficient Large Language Models through Compact Kernelization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.462608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.462608Z digest=sha256:794c39edf91f400f17c735f2389dd559e4f2bb9f5487231b7173cb8e2c1ec2b2

Observation 78264015-2891-49f7-8732-d3f03fdaa226 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:44.602537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:44.602537Z digest=sha256:2a20c06a5172eb577eb690a0ae8dd1a6e1814eaf6dcdea21fd418590bf5bb806

Observation 70cf876f-4c6d-4a79-a797-30c1114bd463 · outbound

This paper cites Ravichandran, A.

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency Ravichandran, A

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:20:46.532723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:20:45.521317Z digest=sha256:a2d1c031d062eafb66f4fa6cdc1ff8efa50af67bc486e7b18f5183f81c509762

Pith citing papers

No inbound Pith citation observations are available.