Pith. sign in

Paper Citation Record · LEDGER

Attention Mechanism, Max-Affine Partition, and Universal Approximation

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2504.19901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19901 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:39.652481Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b507613a-4654-4f7a-a466-b54a2cea6533 · outbound

This paper cites GPT-4 Technical Report.

Attention Mechanism, Max-Affine Partition, and Universal Approximation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.563435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.563435Z digest=sha256:2a2f6c355f68da746f1ac286d803fa3fc5073f6359b24fe41f2db9141464cc00

Observation 2be047a2-6e9d-4a15-85fa-95bf7424e1e5 · outbound

This paper cites Superiority of softmax: Unveiling the performance edge over linear attention.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Superiority of softmax: Unveiling the performance edge over linear attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.592971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.592971Z digest=sha256:43ca8508973cf92f9a380b84419182fdad5b7c5faa2183f060abe2ebfa96b349

Observation 239a4c30-fddb-4f79-a561-9d1bf322860a · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Attention Mechanism, Max-Affine Partition, and Universal Approximation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.597272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.597272Z digest=sha256:f206907bbb7631e15537ef7c30cd689cf06da4f610a9d8f2d9313dcdf0c3c855

Observation 0a54420d-e062-4547-a14d-f7cd99503ed5 · outbound

This paper cites Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.606729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.606729Z digest=sha256:59b314d57f90ac4605388598ae2581ae4d599344c1df8b3048fe3790e0bca9e6

Observation ea2009c0-44ee-47e0-b72b-37f6bbe1b1bf · outbound

This paper cites On the Optimal Memorization Capacity of Transformers.

Attention Mechanism, Max-Affine Partition, and Universal Approximation On the Optimal Memorization Capacity of Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.611166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.611166Z digest=sha256:fd10388a3a3271b939819b61e675275ce5df550b40d0119b62011ca661abf547

Observation d6f2e11b-9e34-421d-9f8c-524bb324f063 · outbound

This paper cites Transformers Provably Solve Parity Efficiently with Chain of Thought.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Transformers Provably Solve Parity Efficiently with Chain of Thought

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.615441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.615441Z digest=sha256:f14ad1a8c685e6ac5dec45d95ee2bf4af2c32cc494502b5ba94d9fbb74f39b92

Observation 7a1922f4-66b7-4a16-afdf-cf9a795ad73f · outbound

This paper cites Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.619918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.619918Z digest=sha256:4fee02c944f912ae9e662d7fac3fea8bd5d86a4fb9fd998ef6f23caea0f44ec9

Observation bde8a35c-b449-439a-88ba-253bf7ebf5f2 · outbound

This paper cites In-context Learning and Induction Heads.

Attention Mechanism, Max-Affine Partition, and Universal Approximation In-context Learning and Induction Heads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.629262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.629262Z digest=sha256:8cc9448fd1d51ac3a96b2381ea1d778448405f4b68a985173f7f1ce12da8b321

Observation 0ba2cba4-d88c-45ee-9920-30b4ca025ece · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Transformers, parallel computation, and logarithmic depth

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.633644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.633644Z digest=sha256:a9b62f5b0328da3cd43858615c8ce6e9f6e3ad9b0adb4508965515fb0405f556

Observation cc945918-f96a-45da-8c6d-81d110d5c562 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Attention Mechanism, Max-Affine Partition, and Universal Approximation LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.638354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.638354Z digest=sha256:50a423279dd92911cf0faff61bf749ce4b61c80486fa23fe98b3d4f3566337ca

Observation 3bfb064f-9405-46f3-af0f-de6310a4f201 · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.647398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.647398Z digest=sha256:c2a297bcd5a20670154a976b0282b0dc975cea8787f77dce29d9a0bdbe830ff1

Observation 1c4f421b-27cb-4979-82af-b5f27f6886f7 · outbound

This paper cites Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.583561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.583561Z digest=sha256:d49540eacc6134afd06e53ff4199d00bf7b226087903440dca5b9e33098139e4

Observation f6e7943d-c231-4bdf-bf24-8c190f154d70 · outbound

This paper cites Unified Training of Universal Time Series Forecasting Transformers.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Unified Training of Universal Time Series Forecasting Transformers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.642809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.642809Z digest=sha256:afc940f8cbae6c59d09d5bc586553a6487b0853a17da103f282d1de60419acb8

Observation 31ee2ee9-d86a-424e-943c-ca46ec15186d · outbound

This paper cites DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome.

Attention Mechanism, Max-Affine Partition, and Universal Approximation DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.652481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.652481Z digest=sha256:f5acddcd1239e0812e358d949a27f443ff82cc6fc9ebc96de2489ae11358b82b

Observation 7c8e4ee2-65d2-4662-a941-dc4737d5bdc8 · outbound

This paper cites Construction of neural nets using the radon transform.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Construction of neural nets using the radon transform

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:49:40.028546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:49:39.578926Z digest=sha256:4aea8afa0fe58b257a133391a203ffc4cbb364b7aaa0eb964852304caa226529

Observation 5fa68262-549b-45e8-9ad2-6af9e69e5ae0 · outbound

This paper cites The Llama 3 Herd of Models.

Attention Mechanism, Max-Affine Partition, and Universal Approximation The Llama 3 Herd of Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.602242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.602242Z digest=sha256:97edc303dece12e8cd67070cdbf5fa3cf25609c4c1bcc7e069c273d8f32f8729

Observation 289511ff-3ec4-438f-be27-4f0afc7f3bad · outbound

This paper cites Memorization Capacity of Multi-Head Attention in Transformers.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Memorization Capacity of Multi-Head Attention in Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.624488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.624488Z digest=sha256:d02c18de4e4ae8a3113b8c6f0d84a2c1c22b04c81991bdbb6f985c307965a065

Observation e6aae1c7-384d-41d5-b36d-978f6a39d766 · outbound

This paper cites Fundamental Limitations on Subquadratic Alternatives to Transformers.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Fundamental Limitations on Subquadratic Alternatives to Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.568945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.568945Z digest=sha256:d85663c1a0236b9c3531c1309ac7d6ae4c22d198ea816b6f25d71e1b3a4cf1e2

Observation 449140ae-de5b-49b3-b0e9-71b2f5e89f4e · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Gireeja Ranade Sastry, Amanda Askell, et al.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Gireeja Ranade Sastry, Amanda Askell, et al

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:49:40.047596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:49:39.573944Z digest=sha256:55b75f219849b0e047d2bb5e113b9e2709149a69e5bba24a0a50bcade05be586

Observation 26894e0e-f0a5-48e1-9329-2652a5012270 · outbound

This paper cites Provably learning a multi-head attention layer.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Provably learning a multi-head attention layer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.588423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.588423Z digest=sha256:992d9a12d12d4d2be89b322b568d9850966a65ecedf82f43b00a9f48f3fd0fcf

Pith citing papers

No inbound Pith citation observations are available.