Pith. sign in

Paper Citation Record · LEDGER

Length Generalization of Causal Transformers without Position Encoding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2404.12224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.12224 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.257094Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c72e8e0f-57f3-4f34-b4ce-a6db9d23ebcd · inbound

Scalable-Softmax Is Superior for Attention cites this paper.

Scalable-Softmax Is Superior for Attention Length Generalization of Causal Transformers without Position Encoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257094Z digest=sha256:184a2f41fc5104ba12bfce6861092cd601c20738be17336e834653d8bfa8df21

Observation 295eda4b-cbf4-4ef2-8810-4abfc8d4b5dc · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Length Generalization of Causal Transformers without Position Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.381466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.381466Z digest=sha256:b031c5d47e2a77be2d6608add05a1c6e778743c29ca68e973d0176d65b74e2d3

Observation a7586f8e-18a6-4212-9f91-f085bd39dd63 · inbound

Understanding Transformer from the Perspective of Associative Memory cites this paper.

Understanding Transformer from the Perspective of Associative Memory Length Generalization of Causal Transformers without Position Encoding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.463146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.463146Z digest=sha256:5b7c1ffff41dd6f82a722c87dfaec9cb70ee8571b3d973492c7b7bdc23740489

Observation a947700b-699d-4ab9-b363-beabfe0252b0 · inbound

Home-made Diffusion Model from Scratch to Hatch cites this paper.

Home-made Diffusion Model from Scratch to Hatch Length Generalization of Causal Transformers without Position Encoding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:34:37.047557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:34:37.047557Z digest=sha256:bce9588efced66d06f64ea6df04d16363c5ecbc689cb04927311ae5e5a3e8935

Observation 48e16509-8bb2-4695-9968-b1ebba7cf5a9 · inbound

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings cites this paper.

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings Length Generalization of Causal Transformers without Position Encoding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:20:51.337939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:50:43.716813Z digest=sha256:d0e06191297c41eba7df9f6dad854b6e376ca89e00151b4b67a57c94b55da12b

Observation e7be42e3-9ee3-4f0e-822c-25ace2c4df9d · inbound

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Length Generalization of Causal Transformers without Position Encoding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.034396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:52:52.985532Z digest=sha256:4c74e63afb05b6ae38a5c752e5ac9eb6eae7dfd8ff3c57b9439174a1935cb3c3

Observation 4b4bbb94-280d-4f19-84a1-fbee2b9f7a9b · inbound

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Length Generalization of Causal Transformers without Position Encoding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.720777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:38:43.250420Z digest=sha256:7e0e85122082a761080a80325e899c47cbdf38c1f3bb206e89f8f4499765e6fc