Pith. sign in

Paper Citation Record · LEDGER

Transformer learns the cross-task prior and regularization for in-context learning

As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2505.12138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12138 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:46:23.953496Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:00:10.889049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:18:14.285893Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d19dcb9-a133-46d0-8acb-0664486afdf2 · outbound

This paper cites Learning regularization parameters of inverse problems via deep neural networks.Inverse Problems, 37(10):105017, sep 2021.

Transformer learns the cross-task prior and regularization for in-context learning Learning regularization parameters of inverse problems via deep neural networks.Inverse Problems, 37(10):105017, sep 2021

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.710771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.751405Z digest=sha256:b9c678a4a15307eedcb2bb664eba2ecc261a84eb0b807fefc4f94d31567e68d0

Observation 9f4cfecf-ecd5-4407-b985-6554f7187278 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.Advances in Neural Information Processing Systems, 36:45614–45650, 2023.

Transformer learns the cross-task prior and regularization for in-context learning Transformers learn to implement preconditioned gradient descent for in-context learning.Advances in Neural Information Processing Systems, 36:45614–45650, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.694243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.757753Z digest=sha256:9c830febb0e223f53ea17fa8acadcc74ff0fe029a0e4005b24242f5f1797895e

Observation 3ed22182-68c5-4f3e-8d93-92542caacb88 · outbound

This paper cites Transformers can learn to solve linear-inverse problems in-context.

Transformer learns the cross-task prior and regularization for in-context learning Transformers can learn to solve linear-inverse problems in-context

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.678249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.762840Z digest=sha256:f41c48da7fe145d6b6ed2efc7def8f06dff6096b7f15f5b0c498abe6105fd558

Observation dcddeb08-b33f-40d5-bf8e-fd18f179e202 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Transformer learns the cross-task prior and regularization for in-context learning What learning algorithm is in-context learning? investigations with linear models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.662272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.769030Z digest=sha256:d7cb10553feae8486d29bc5639f06ad5aa0418e36262e5abe43f67dfcc2dfefc

Observation 397f7f3a-6fbf-48c6-8a2c-b756664f83e9 · outbound

This paper cites Bayesian scaling laws for in-context learning.arXiv preprint arXiv:2410.16531, 2024.

Transformer learns the cross-task prior and regularization for in-context learning Bayesian scaling laws for in-context learning.arXiv preprint arXiv:2410.16531, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.774041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.774041Z digest=sha256:c04dc436613bc52b642ba4d002d554118f18f33b340bd69fdf28e0c58194d4f2

Observation 69764f90-86a7-4942-8ddf-fea02284e3c4 · outbound

This paper cites Transformers as statisticians: Provable in-context learning with in-context algorithm selection.Advances in neural information processing systems, 36:57125–57211, 2023.

Transformer learns the cross-task prior and regularization for in-context learning Transformers as statisticians: Provable in-context learning with in-context algorithm selection.Advances in neural information processing systems, 36:57125–57211, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.779394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.779394Z digest=sha256:08bdca2e4b99a200f0ac51d11b7f6abaa7f58fa5ebe00be5634d61ba93df40be

Observation bb0c60f4-1bd6-4caa-909a-5096606d55c7 · outbound

This paper cites Understanding in-context learning in transformers and llms by learning to learn discrete functions.

Transformer learns the cross-task prior and regularization for in-context learning Understanding in-context learning in transformers and llms by learning to learn discrete functions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.633770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.785345Z digest=sha256:6ae23cd8487fae47d7232cbeb81f4eae1e2bbfce6578a62445d71cdb7dddf021

Observation df1d17df-e4b7-4788-a032-ffa776a1c313 · outbound

This paper cites Language Models are Few-Shot Learners.

Transformer learns the cross-task prior and regularization for in-context learning Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.790350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.790350Z digest=sha256:61f8e45056d07d48ee6c9241160879c94a5508c00fcd1aba69bc1528ba288f4c

Observation 9fa55b68-17a2-4641-9288-5f85704f3939 · outbound

This paper cites Choose a transformer: Fourier or Galerkin.Advances in neural information processing systems, 34:24924–24940, 2021.

Transformer learns the cross-task prior and regularization for in-context learning Choose a transformer: Fourier or Galerkin.Advances in neural information processing systems, 34:24924–24940, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.617309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.796703Z digest=sha256:fdfe646401ef6ce3ac54d841b33b58002d908bf43bbe6ee8b7651cd52b92a63d

Observation 0d72db8b-aa30-4e5c-b87a-0d3964b0dd57 · outbound

This paper cites Deformable cross-attention transformer for medical image registration.

Transformer learns the cross-task prior and regularization for in-context learning Deformable cross-attention transformer for medical image registration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.600376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.802886Z digest=sha256:933f2ec253516929453984f3e632a0510ee7c34a22550222e6cb90181397ef78

Observation d33fb7fd-9f34-47e5-ad60-9455c6f727b7 · outbound

This paper cites In-context learning with transformers: Softmax attention adapts to function lipschitzness.Advances in Neural Information Processing Systems, 37:92638–92696, 2024.

Transformer learns the cross-task prior and regularization for in-context learning In-context learning with transformers: Softmax attention adapts to function lipschitzness.Advances in Neural Information Processing Systems, 37:92638–92696, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.584414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.808025Z digest=sha256:01352240bf30bbec46552ab1f96dafde5ea056f69546245fd1edfe99a3c1704e

Observation b1dec4f6-173a-4fb5-b4da-584aa372bda9 · outbound

This paper cites Why can GPT learn in-context? Language models implicitly perform gradient descent as meta-optimizers.

Transformer learns the cross-task prior and regularization for in-context learning Why can GPT learn in-context? Language models implicitly perform gradient descent as meta-optimizers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.568983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.813146Z digest=sha256:1c7614c34c39f6c9a92341334d458c4f0da63e5d2a4a72f02a2c39565fc20f95

Observation 82a1d72e-1ca2-403f-98b9-88e4199db50b · outbound

This paper cites an unresolved cited work.

Transformer learns the cross-task prior and regularization for in-context learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:46:24.552266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.817780Z digest=sha256:7bd5e90ea6362334de320b631ddbf73aeb2b2135b48848e729ff438ba4579c6b

Observation 98940d5d-c17d-42da-8909-f114b1aaa33c · outbound

This paper cites To be or not to be stable, that is the question: understanding neural networks for inverse problems.SIAM Journal on Scientific Computing, 47(1):C77–C99, 2025.

Transformer learns the cross-task prior and regularization for in-context learning To be or not to be stable, that is the question: understanding neural networks for inverse problems.SIAM Journal on Scientific Computing, 47(1):C77–C99, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.526941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.827463Z digest=sha256:914f6d6b7b99bf400a0b859393e887a5072c31cf13b00dfc546cc3ce0de310e3

Observation 7b89f89b-759e-44a4-9bf7-d3ccecd894c0 · outbound

This paper cites Ambiguity in solving imaging inverse problems with deep-learning-based operators.Journal of Imaging, 9(7):133, 2023.

Transformer learns the cross-task prior and regularization for in-context learning Ambiguity in solving imaging inverse problems with deep-learning-based operators.Journal of Imaging, 9(7):133, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.510460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.832227Z digest=sha256:1e0ea3cded37189593efc97e6a1af156a2abfe5268a51f1338761cc41bf8adf5

Observation 8da6f98c-b49c-4d88-ae5a-b274e153dd03 · outbound

This paper cites Transformers are Universal In-context Learners.

Transformer learns the cross-task prior and regularization for in-context learning Transformers are Universal In-context Learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.837321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.837321Z digest=sha256:f5ac0efd5d1702ff72240ac49fe892664aafa49d8846bd8e1cfd945c0dec6c61

Observation 975bffda-da12-47cc-89ef-8187fae0fcb5 · outbound

This paper cites What can transformers learn in- context? a case study of simple function classes.Advances in Neural Information Processing Systems, 35:30583–30598, 2022.

Transformer learns the cross-task prior and regularization for in-context learning What can transformers learn in- context? a case study of simple function classes.Advances in Neural Information Processing Systems, 35:30583–30598, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.494786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.842381Z digest=sha256:e56ffebec87da0f6048e241c54b287edb03ade29c2c302615d40d0315176206f

Observation 0be534e9-d177-40d3-8ba7-b86be3f1c509 · outbound

This paper cites Generalized cross-validation as a method for choosing a good ridge parameter.Technometrics, 21(2):215–223, 1979.

Transformer learns the cross-task prior and regularization for in-context learning Generalized cross-validation as a method for choosing a good ridge parameter.Technometrics, 21(2):215–223, 1979

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.847382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.847382Z digest=sha256:d6cc5a81aca236a18e90e13f2357aecf9e44442456bb21a473c12e12f1c16bbb

Observation 11758794-f318-4b39-817b-db40b70e5a34 · outbound

This paper cites Transformer meets boundary value inverse problems.

Transformer learns the cross-task prior and regularization for in-context learning Transformer meets boundary value inverse problems

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.467267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.853378Z digest=sha256:b683a1984dc56d153b807336071a9bceb241737b52dce5b4dfd5c28c99a6a636

Observation 4912b199-ffe8-440f-baeb-5f37346f33fa · outbound

This paper cites SIAM, 1998.

Transformer learns the cross-task prior and regularization for in-context learning SIAM, 1998

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.449896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.857775Z digest=sha256:0e031e4995b1110f6dc863cf6c74bc690320b7997aa23d3199ffd78f45cadb9a

Observation a06de703-9998-45a5-888e-1b006ae1dd72 · outbound

This paper cites The L-curve and its use in the numerical treatment of inverse problems.

Transformer learns the cross-task prior and regularization for in-context learning The L-curve and its use in the numerical treatment of inverse problems

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.432558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.862607Z digest=sha256:8781b4505bb8b5d79fc50aa4476bbdaa832c721f5cfeb53cd8b260590421a3d5

Observation 79d8981b-91ac-43be-9fe5-4afcfa0bcd39 · outbound

This paper cites an unresolved cited work.

Transformer learns the cross-task prior and regularization for in-context learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:46:24.410836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.867869Z digest=sha256:05ef7352b59cc7021d875170937d4a90a466930b64b6fc41c151dc5bdacaa7c5

Observation 7b28ddae-f7f3-4d4a-94ef-903fa5917723 · outbound

This paper cites Hoerl and Robert W.

Transformer learns the cross-task prior and regularization for in-context learning Hoerl and Robert W

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.873684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.873684Z digest=sha256:39f2efe132d0f733529ffe7f4b14ce2ab6900d947cd7bf3de4a97a983ba88e1b

Observation edc314a8-051a-4759-9ec6-2439127adb03 · outbound

This paper cites Transformers as algorithms: Generalization and stability in in-context learning.

Transformer learns the cross-task prior and regularization for in-context learning Transformers as algorithms: Generalization and stability in in-context learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.377738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.878411Z digest=sha256:90ec57d3660256c3fd16be25e0e3a3a9f318b9775cc12c97539d17132ce87b46

Observation a979edb6-4104-4664-a23a-442991ed21af · outbound

This paper cites Asymptotic theory of in-context learning by linear attention.arXiv preprint arXiv:2405.11751, 2024.

Transformer learns the cross-task prior and regularization for in-context learning Asymptotic theory of in-context learning by linear attention.arXiv preprint arXiv:2405.11751, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.883234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.883234Z digest=sha256:86bbd71d5bb58e2ffdbcd334032c71ff9bb758545c678a9068299686cc7cc299

Observation 41776518-b1fe-42f8-901b-4b5e2c90d6f0 · outbound

This paper cites One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention.

Transformer learns the cross-task prior and regularization for in-context learning One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.360276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.887758Z digest=sha256:4c22eb2898223493aef3cb0d099628c1ed7d1f5f91f6c8bfe613606eb8b17881

Observation 149ffed9-0316-4374-be07-cdd44c46cdc2 · outbound

This paper cites Pretrained transformer efficiently learns low- dimensional target functions in-context.Advances in Neural Information Processing Systems, 37:77316– 77365, 2024.

Transformer learns the cross-task prior and regularization for in-context learning Pretrained transformer efficiently learns low- dimensional target functions in-context.Advances in Neural Information Processing Systems, 37:77316– 77365, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.342839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.892457Z digest=sha256:7acaefcddd65e47efe7c73b526b31c338a5bd681f4755e0e73eacf4bd6d0eec5

Observation 442bccf2-fefd-4b1f-94b6-97cfa513356f · outbound

This paper cites Vito: Vision transformer-operator.Computer Methods in Applied Mechanics and Engineering, 428:117109, 2024.

Transformer learns the cross-task prior and regularization for in-context learning Vito: Vision transformer-operator.Computer Methods in Applied Mechanics and Engineering, 428:117109, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.897044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.897044Z digest=sha256:ca10aba5f39f6d10e66635228a6cb9adaea1490edd152766f94953970d4817e3

Observation 1aae1460-0c73-4353-9935-039595398479 · outbound

This paper cites Transformers can optimally learn regression mixture models.

Transformer learns the cross-task prior and regularization for in-context learning Transformers can optimally learn regression mixture models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.315511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.901684Z digest=sha256:794ba80b5f1468df8b6116b9a5f2e15f0a348c7f8007e4d48aad14a30cfff932

Observation eca3338d-e2b0-4b9c-8b5d-f1a08a0c82fe · outbound

This paper cites The mechanistic basis of data dependence and abrupt learning in an in-context classification task.

Transformer learns the cross-task prior and regularization for in-context learning The mechanistic basis of data dependence and abrupt learning in an in-context classification task

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.297286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.906383Z digest=sha256:44cbd07ac37dde43159582059b1ae67591e06b14404fd0ffa7625d985a5c1b29

Observation f3b080db-89c4-418e-bbb1-9dc1e121f1ea · outbound

This paper cites JoMA: Demystifying multilayer transformers via joint dynamics of MLP and attention.

Transformer learns the cross-task prior and regularization for in-context learning JoMA: Demystifying multilayer transformers via joint dynamics of MLP and attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.280331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.910994Z digest=sha256:ade23fd4510a09c4a758d71cacc77358a1ad12d9b524f9151ad454bf965e4706

Observation 096b5811-1950-4d39-b0db-f026fdc8ebca · outbound

This paper cites Cambridge University Press, 2018.

Transformer learns the cross-task prior and regularization for in-context learning Cambridge University Press, 2018

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.915694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.915694Z digest=sha256:b094f84d2fd49c0b4cea559c4c15809c59b9093c970f7d33f230054f31683546

Observation 96be4e06-939a-4342-b483-978156581da5 · outbound

This paper cites Linear Transformers are Versatile In-Context Learners.

Transformer learns the cross-task prior and regularization for in-context learning Linear Transformers are Versatile In-Context Learners

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:46:24.000234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.920567Z digest=sha256:fc94bef78a0094f4ef3d63dc67b4624de6828d892a5c908fa51209c9544e1413

Observation a15ab009-6ead-4ab6-9440-8e9cc20d3487 · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformer learns the cross-task prior and regularization for in-context learning Transformers learn in-context by gradient descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.926413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.926413Z digest=sha256:e3d8c5a9eef16b56be86fec8f5dfa737afa95e68ee2f4e6e1aa35f9791f15bb8

Observation 41ab8d34-6192-4cb4-a35c-09ad8c037295 · outbound

This paper cites The learnability of in-context learning.Advances in Neural Information Processing Systems, 36:36637–36651, 2023.

Transformer learns the cross-task prior and regularization for in-context learning The learnability of in-context learning.Advances in Neural Information Processing Systems, 36:36637–36651, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.243059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.932280Z digest=sha256:10204dc344f15529f803f57e368fdef0d99e3f603755ed66c23c1cd80e9c0f57

Observation 820eea7e-8f09-4071-9221-69eea80c65a7 · outbound

This paper cites An explanation of in-context learning as implicit bayesian inference.

Transformer learns the cross-task prior and regularization for in-context learning An explanation of in-context learning as implicit bayesian inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.937235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.937235Z digest=sha256:8b181fba9365121eaf3e06db72d4f688f88e88120e62b6f262af36c4fa50aec0

Observation 13a6ac49-2d0a-4b56-858b-dcdf0b5e15e5 · outbound

This paper cites A data-driven peridynamic continuum model for upscaling molecular dynamics.Computer Methods in Applied Mechanics and Engineering, 389:114400, 2022.

Transformer learns the cross-task prior and regularization for in-context learning A data-driven peridynamic continuum model for upscaling molecular dynamics.Computer Methods in Applied Mechanics and Engineering, 389:114400, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.942436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.942436Z digest=sha256:1fb01de706fabfd8297cd8c5732bb09e8e1601097233fcf1d53f5de0f2cff998

Observation c517467e-300b-4814-bedb-ec326c72759f · outbound

This paper cites Nonlocal attention operator: Materializing hidden knowledge towards interpretable physics discovery.

Transformer learns the cross-task prior and regularization for in-context learning Nonlocal attention operator: Materializing hidden knowledge towards interpretable physics discovery

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.204553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.948529Z digest=sha256:12663a0014fb8a225a47964afd3d62fcdd7ec633fd7ef29903859f1b90d99654

Observation d67645de-80f4-4794-9983-3c8e4c048595 · outbound

This paper cites Trained transformers learn linear models in-context.

Transformer learns the cross-task prior and regularization for in-context learning Trained transformers learn linear models in-context

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:24.186326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:46:23.953496Z digest=sha256:ba353a4411e3aec0662c09d6b15cf8062ae209a32d0ceefc4866dcbf10af943d

Observation 6db64c5b-a853-4a2e-8464-a8aaaa1e9005 · outbound

This paper cites an unresolved cited work.

Transformer learns the cross-task prior and regularization for in-context learning Unresolved cited work

Reference 375

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:23.822587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:23.822587Z digest=sha256:f9fd1379ecb97aee3c250de0200b83b0a709c7aa5d159733498b58cc8de8a2fd

Pith citing papers

Observation caa14c7f-ac78-4130-ad7a-729c87e48da2 · inbound

Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics Discovery cites this paper.

Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics Discovery Transformer learns the cross-task prior and regularization for in-context learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:03.142987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:03.142987Z digest=sha256:0707343fb65545fd197d57e71c062961203d81df98b17dfc76c368df5eee9df4

Observation b30d1b7e-8e02-4d7d-bbe1-79062ff5931a · inbound

An Attention-based Spatio-Temporal Neural Operator for Evolving Physics cites this paper.

An Attention-based Spatio-Temporal Neural Operator for Evolving Physics Transformer learns the cross-task prior and regularization for in-context learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:18:14.289000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:18:14.110736Z digest=sha256:63f2beb277a17a947438b2b16bd9a325a692a1020e51c0a78558a974694555c8

Observation 12ce7a86-dad9-4ff2-9306-86d297565572 · inbound

Learning Causal Graphs at Scale: A Foundation Model Approach cites this paper.

Learning Causal Graphs at Scale: A Foundation Model Approach Transformer learns the cross-task prior and regularization for in-context learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:10.889049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:10.889049Z digest=sha256:be8874400a8edcfba727f09d177b01331ae5f784c7f218950d98de476487f10e