Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't In-Context Learn Least Squares Regression

As of 17 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.09440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09440 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:14.974580Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b23e4a6-bdd4-4d8f-91da-b93f14f117a4 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Transformers Don't In-Context Learn Least Squares Regression What learning algorithm is in-context learning? investigations with linear models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.202146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.923756Z digest=sha256:eb22309c70aabd04efb12c051ae35e2224fdc38d071a44d4c382853717a43ac1

Observation a1ad1156-3c84-4937-8e68-f6da49dfa4f8 · outbound

This paper cites Bayesian scaling laws for in-context learning.

Transformers Don't In-Context Learn Least Squares Regression Bayesian scaling laws for in-context learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.926769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.926769Z digest=sha256:679f7dd6c69a8fcf3e8494808a9e7cf4f28c680ae1ba067629b9abf2ef5694ee

Observation 11ab1bdb-0139-4697-a754-2ce8921c8c86 · outbound

This paper cites Learning theory from first principles.

Transformers Don't In-Context Learn Least Squares Regression Learning theory from first principles

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.929310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.929310Z digest=sha256:df2b5275f390eb6e28779803cdf2017e6b4104f2f2f7ea7e47e6db82f87dad0b

Observation 254c2b31-7047-4fe2-bc0e-72c09099767b · outbound

This paper cites Training with noise is equivalent to tikhonov regularization.

Transformers Don't In-Context Learn Least Squares Regression Training with noise is equivalent to tikhonov regularization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.189149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.932028Z digest=sha256:46a21c26be305f8713c0f083b301150238a9156d911c40e90fd015a30ecee8ea

Observation bc88ab7d-e034-4be9-8fd0-2f3599295251 · outbound

This paper cites Language Models are Few-Shot Learners.

Transformers Don't In-Context Learn Least Squares Regression Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.934798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.934798Z digest=sha256:7a78cddc285e4460e42668adc5bd92b5ef630e66b3142ce49763530261bc8dfb

Observation 0542686e-d2cb-49c4-ac0a-f107f01ae275 · outbound

This paper cites A survey on in-context learning.

Transformers Don't In-Context Learn Least Squares Regression A survey on in-context learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.181492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.937609Z digest=sha256:3ad714e6f9822b79d455d45281e2133e902e674eef27c6e82883c2a6aa035346

Observation b331e331-95f0-47ef-9d6b-58dd437b793f · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't In-Context Learn Least Squares Regression A mathematical framework for transformer circuits

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.940846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.940846Z digest=sha256:26a05e3be7a737ced7404619a1b5cd9ca7a9085af8cd68855e4a036796f96e39

Observation 949a9976-c58b-4668-aba9-159d7e45e536 · outbound

This paper cites What Can Transformers Learn In-Context? A Case Study of Simple Function Classes.

Transformers Don't In-Context Learn Least Squares Regression What Can Transformers Learn In-Context? A Case Study of Simple Function Classes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.943494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.943494Z digest=sha256:0f4edc1d567885893e9e69e36e7217e362bd6bc395aadf4a8311d6e79a872a7d

Observation 4dd223b8-b3d4-405d-9892-b5f6c3e7dcd8 · outbound

This paper cites Smith, Vudtiwat Ngampruetikorn, and David J.

Transformers Don't In-Context Learn Least Squares Regression Smith, Vudtiwat Ngampruetikorn, and David J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.168426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.946522Z digest=sha256:b5eaf2a02e39601a1bff9f01a486a1c0bb8fb58292c61fc2fc155e53ad9ad82c

Observation cfb30f18-ab1e-4313-980c-4a1bf410b6fa · outbound

This paper cites Cauchy's method of minimization.

Transformers Don't In-Context Learn Least Squares Regression Cauchy's method of minimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.160373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.948848Z digest=sha256:48b026a539b5fb8442b97020828698f6691ee020d7534726f7027f2733ccbebc

Observation 2ae55cb5-ad13-4240-98d7-d4f47519fe95 · outbound

This paper cites Understanding Catastrophic Forgetting in Language Models via Implicit Inference.

Transformers Don't In-Context Learn Least Squares Regression Understanding Catastrophic Forgetting in Language Models via Implicit Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.951300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.951300Z digest=sha256:fd7966c6f1939eb19d10f6c90777b477e2d8233be42d644d1cff85f303babaeb

Observation 35b6bd68-56f0-489a-be01-bf93c7a2308a · outbound

This paper cites Predictive multiplicity in classification.

Transformers Don't In-Context Learn Least Squares Regression Predictive multiplicity in classification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.152746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.954133Z digest=sha256:470c2e20ac266844231d18287680a9a0f91fa571b5c32be2b01cd82f41469377

Observation 51c9e342-52d9-4a2a-8e7d-48d5f35d2ecc · outbound

This paper cites Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve.

Transformers Don't In-Context Learn Least Squares Regression Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.956649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.956649Z digest=sha256:f25084b0791915553ecc93937b11240a46d27b45b535114b05b7e85c58229c7f

Observation 1e99849d-ecff-486f-b090-8dd7783026df · outbound

This paper cites More Data Can Hurt for Linear Regression: Sample-wise Double Descent.

Transformers Don't In-Context Learn Least Squares Regression More Data Can Hurt for Linear Regression: Sample-wise Double Descent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.959393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.959393Z digest=sha256:d70ec05e4f655203f3e62a98381f82d2292071bcab2ce8b6f73948d7bb92cb5a

Observation c70fea0b-1245-46dc-9998-631bc4476c0c · outbound

This paper cites In-context Learning and Induction Heads.

Transformers Don't In-Context Learn Least Squares Regression In-context Learning and Induction Heads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.961955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.961955Z digest=sha256:3762982056989ac10b4a0b21d3d4f97b8aecd04045346b1179e84963689c02d0

Observation f407d3b4-0610-48e5-935a-fecbafd129aa · outbound

This paper cites Pretraining task diversity and the emergence of non-bayesian in-context learning for regression.

Transformers Don't In-Context Learn Least Squares Regression Pretraining task diversity and the emergence of non-bayesian in-context learning for regression

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.144127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.964504Z digest=sha256:bc7975d6f149c3c6c7cbbaf81c3bfdd635c812000c6e99d2f3e12b7a31403579

Observation 4637800a-69f2-4bdf-a36b-dcc194d23490 · outbound

This paper cites Do pretrained Transformers Learn In-Context by Gradient Descent?.

Transformers Don't In-Context Learn Least Squares Regression Do pretrained Transformers Learn In-Context by Gradient Descent?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.967182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.967182Z digest=sha256:72febb0bab6753abf23092669214f5a1d11c1a6e14dfb326a229e5693a90b014

Observation 3645c945-832d-44af-b92f-5d82f0d7e7aa · outbound

This paper cites Attention Is All You Need.

Transformers Don't In-Context Learn Least Squares Regression Attention Is All You Need

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.969689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.969689Z digest=sha256:a2aa0e2bfe77303e837ec8e92fc4ca4179b65aebf5c1476e6c1090b12f15b853

Observation a89c9943-b8c2-42d6-807a-a9a287c3f86a · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformers Don't In-Context Learn Least Squares Regression Transformers learn in-context by gradient descent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.972174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.972174Z digest=sha256:1f895c10c038dc1bc42bb111bf6422d1077b63f71f0f10ba62287f0f730fac0b

Observation f6148ffc-202a-4a66-8b35-579ec08368da · outbound

This paper cites Qwen3 Technical Report.

Transformers Don't In-Context Learn Least Squares Regression Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.974580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.974580Z digest=sha256:6b2dc626fda1821d9bd58d77b938a2a10885fdd6ce82c28767e90de9ea71b1fd

Pith citing papers

No inbound Pith citation observations are available.