Pith. sign in

Paper Citation Record · LEDGER

Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2312.06528.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06528 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:17.138258Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T07:34:02.998394Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.138258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.138258Z digest=sha256:bf8ab7dda1a1436a4623e275da0289a2bc4891b0314b115ded02682c63a6d1d5

Observation b9de054c-1af0-4782-a4fd-5f0fc9b868b9 · inbound

Spectral Transformer Neural Processes cites this paper.

Spectral Transformer Neural Processes Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:34.735349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:04:14.880310Z digest=sha256:4ce5cc244206e551220b1d1703d5994e2ed602c3672ec3615956f04c3e4b0392

Observation d94cfa03-61e1-4974-8dd2-0104b67b1879 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:32.087928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:6ab6a21dc9409239b2b8299586ba4eb497597e820e7de7a8be00f592ae4b93e4

Observation eaefecc6-06a6-4095-9039-12c0374c2ae9 · inbound

Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity cites this paper.

Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:34:02.999865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-21T07:29:50.058460Z digest=sha256:a3b13c0974f4ea3e0888af66c6f1fe375c56f37d0ebb6d55843ada31516d827e

Observation e5ea04a7-25a0-4d51-a9d4-a1c7108f0e37 · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 161

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:03.096003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:03.096003Z digest=sha256:6d69a073c35352515912d7c2bc68054e4da0feae547f4caada8af3c3c84ca66e