Pith. sign in

Paper Citation Record · LEDGER

Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2408.10189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.10189 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:48.707551Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65704c0f-3336-4321-920c-eff8fd961f2d · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.570139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:4aab01d9020bdb1028e5a518c7645a40f35e2d9ee10383075dda2a29a0585a59

Observation eaa20003-4262-4252-8800-37325f79b0d9 · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.707551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.707551Z digest=sha256:b5059f36b3fcd11152c094d9d00e886132dc7ac48b1564cfa959f2aa17b37e34

Observation 8a370696-080e-42f2-8982-19a292cbcad6 · inbound

Maximally-Informative Retrieval for State Space Model Generation cites this paper.

Maximally-Informative Retrieval for State Space Model Generation Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:44.532806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:44.532806Z digest=sha256:7d6a4328b38af75033edad04d20a083c720d3ef866116aebcd78ef689348750e

Observation b841ec1b-1519-4cac-b542-df5ecae1114b · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:01:46.143143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T18:56:48.722344Z digest=sha256:2eebf78b659db6734961fcbeaf56a8f9b0e9d5716a26683bac1d73e6d6f2e4a0

Observation ef2d7010-779a-4abe-9366-74d294dd198a · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:17.013835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:25:17.013835Z digest=sha256:ed9aeb3a70bfb4961edef81ae8139fd51fa0c3c04a071fb61bc94184f5e54c59

Observation 4e606c1f-908a-44d2-bca7-b8731ba14903 · inbound

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling cites this paper.

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:01:10.712163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:39:37.485602Z digest=sha256:745358722298029fc9adc60f137313ef9fab775801c59560983b58b9d65be9f5

Observation d7de1e62-b70d-484a-aad5-a84033a9a1be · inbound

The Transformer as a Polar State Estimator cites this paper.

The Transformer as a Polar State Estimator Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:07.451379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T00:58:28.483037Z digest=sha256:bbd60625e9509d1371c88f06d71a8b67bc33e788033baa03be0a4f645d6179a9

Observation a371e693-b752-48d1-841b-ce50f9acd25d · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.573799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:027197b6dbe3f0ea285a7f89b7f3f42e80fcda4d2d43e7cb13d1cd0e44d00c1f

Observation bb3ec78b-a792-4609-b173-839a757ec85f · inbound

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement cites this paper.

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T18:19:40.899238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:15:21.577196Z digest=sha256:eada1c6dfa0f4a9edf8ced972523e60a24b2a760942d7a3008226c8a862427a6

Observation 281c5cc1-65e6-42b7-aafb-080ff0dc5acf · inbound

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement cites this paper.

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:25:27.954328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:25:17.849729Z digest=sha256:49a8ac6428f61e8ebd927f66db964f0881209a656d46ed2910fc00a8a111bb88

Observation df116046-32ee-4e70-b413-710791d8125e · inbound

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement cites this paper.

PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:07:25.671848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T22:00:40.389728Z digest=sha256:a324ef85cc88074d4c11eaf764f106b6f80af276070598197f9886ded3bac5a2