Pith. sign in

Paper Citation Record · LEDGER

Learning Compositional Functions with Transformers from Easy-to-Hard Data

As of 14 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 11 inbound Pith citation observations for arXiv:2505.23683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23683 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:40.614748Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:15.603859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:33:30.598717Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c791e2bc-b13d-47aa-97a8-afcc6650bf6b · outbound

This paper cites The staircase property: How hierarchical structure can guide deep learning.Advances in Neural Information Processing Systems, 34:26989–27002, 2021.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The staircase property: How hierarchical structure can guide deep learning.Advances in Neural Information Processing Systems, 34:26989–27002, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:32.494326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:32.494326Z digest=sha256:e302b1de3082391b3123ce0402d62575e57caa77895ba5b1e0d838571620a051

Observation ddd8ead5-d8c7-406c-80cd-4067f3f7b8da · outbound

This paper cites The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.442515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:32.596789Z digest=sha256:7568882c4c505463a4b0c079ba7880e850d7639aa04264c82182e13f1cf305a5

Observation a03d8f2f-fec3-480a-9a4c-524dcb10114c · outbound

This paper cites Provable advantage of curriculum learning on parity targets with mixed inputs.Advances in Neural Information Processing Systems, 36:24291–24321, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Provable advantage of curriculum learning on parity targets with mixed inputs.Advances in Neural Information Processing Systems, 36:24291–24321, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:32.708882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:32.708882Z digest=sha256:d2766ba4ecb4d4939b9d177e17a3d89afd90933cbac7d277c9b40d96c416d9dd

Observation 831013f1-6797-4024-ab3a-b34907bde771 · outbound

This paper cites How far can transformers reason? The locality barrier and inductive scratchpad.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How far can transformers reason? The locality barrier and inductive scratchpad

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.275011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:32.865237Z digest=sha256:06270adaa2d0484fa85b179821ef5413dc8bc616e918aacee43a3a67a96e8196

Observation 773f1920-9aa9-4517-95ed-c7f90f9bde19 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.160610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:32.939839Z digest=sha256:0eeade68ef68df9f85c5401104b2421f702be713e3710c279ec4a0a60053fb35

Observation c86c7eff-538c-4a24-8d82-21a2ad229fac · outbound

This paper cites Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.033495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:33.040341Z digest=sha256:48c9d28f2726dde7b4c1ae623ad0e6efdad5d47a1cd74b3f8c887400f669c5fd

Observation 4ab05c7b-4394-466c-b552-cbf39e6178de · outbound

This paper cites The pitfalls of next-token prediction.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The pitfalls of next-token prediction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.870785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:33.173794Z digest=sha256:ca71646d8222b03cc24cbc2a9295ccbca59d8188a4879955798776e49d7a8bf8

Observation b6ed973d-856d-422d-89bd-59af7f89cea5 · outbound

This paper cites Curriculum learning.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Curriculum learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.334050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.334050Z digest=sha256:2f104dc1243020dd87379e0d6e2f96c43ffd98378d86c76bd2d2b06235b2a9f0

Observation 36d3242c-1566-4cde-b737-5d22627eb4f5 · outbound

This paper cites Separations in the Representational Capabilities of Transformers and Recurrent Architectures.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Separations in the Representational Capabilities of Transformers and Recurrent Architectures

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:46:42.789421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:33.485218Z digest=sha256:0f78d4979599b1fda1d58738cb3c27ead6547528e3668c12fbbbc940f41d45da

Observation ccb8cf36-5b8b-4f33-a4be-6228c304285c · outbound

This paper cites Birth of a transformer: A memory viewpoint.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Birth of a transformer: A memory viewpoint

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.760483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:33.649966Z digest=sha256:470b3dbe019ac0d7c6fac28602bfd26f4aaad2415fbcc2007bd27fa01e3aa6c4

Observation c834b064-facb-4bf2-9fb7-c7c5926c5ae1 · outbound

This paper cites Data distributional properties drive emergent in-context learning in transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Data distributional properties drive emergent in-context learning in transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.608752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:33.769852Z digest=sha256:72de9fa9301fbb8d46e7250b24cc05741f1026b81343a3799c23d79a89d05ef9

Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · outbound

This paper cites Theoretical limitations of multi-layer Transformer.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.978754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.978754Z digest=sha256:204e6b826ea53e67fb2ea10b95b54b720a40e7d40f9624c3c7d0485ea2a7fe7c

Observation a597f3a7-8913-43cb-bfa4-167279020fab · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.110434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.110434Z digest=sha256:26331ba4018d0d3df3b478b197990f0cae8c7a6969efffdd1994e15fedaf107a

Observation b134bc0c-7e37-4fb7-824d-df2c5462488e · outbound

This paper cites Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.250512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.250512Z digest=sha256:2ad01531cc49884e3f4af96b8cf9131775917a20c5db524b36d22d5fa6a9fa31

Observation 3ed040cd-8986-4385-b60e-f88d2bf4c500 · outbound

This paper cites Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.394535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.394535Z digest=sha256:ea2edad8235cf3f5c8aa684e899168849583761a58fddab25738deace52c6801

Observation 8bde2b42-80a5-4895-a919-a11e66fefed9 · outbound

This paper cites Neural networks can learn represen- tations with gradient descent.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Neural networks can learn represen- tations with gradient descent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.489754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.489754Z digest=sha256:a9fb64f7624d3994bf4d973e67d73be3920824b2e1dc85fbcead6576b37eb5f2

Observation b5054e96-626f-42fc-be1a-26f48263814f · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Learning Compositional Functions with Transformers from Easy-to-Hard Data From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.662611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.662611Z digest=sha256:0fb737d4e2093e6ac19a598e24c722d259c30ebdfe76a91b22e7664ab7651f0b

Observation 9bfa3d55-b086-4846-a561-d6cee1fbf8ce · outbound

This paper cites The Llama 3 Herd of Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.805557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.805557Z digest=sha256:b11c5e6d948da7d177044c3bca9c6359cd6f44f4082515eef185d16af5f808db

Observation c841ccb3-8b76-48cc-a404-53bbb7898b27 · outbound

This paper cites Faith and Fate: Limits of Transformers on Compositionality.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Faith and Fate: Limits of Transformers on Compositionality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.917688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.917688Z digest=sha256:90ee39399efa3eeb3c015cf31f34fe62331b26f8566b70a28ee22045e866ab95

Observation e4bc1d60-d879-4948-895a-d503ea825cc7 · outbound

This paper cites Learning and development in neural networks: The importance of starting small.Cognition, 48(1):71–99, 1993.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and development in neural networks: The importance of starting small.Cognition, 48(1):71–99, 1993

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.061289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.061289Z digest=sha256:628be2d7659f2b4f7d9a973dcbf879b5743ddf50cdb0771f3c5b293cbc203f15

Observation b543bd7a-7bda-4750-b189-27547524ec73 · outbound

This paper cites Towards revealing the mystery behind chain of thought: a theoretical perspective.Advances in Neural Information Processing Systems, 36, 2024.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards revealing the mystery behind chain of thought: a theoretical perspective.Advances in Neural Information Processing Systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.382613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:35.155644Z digest=sha256:70a4cd8f9e77ab309977a174237f7700f7f03a8466ae256e44843cfe7ae84f62

Observation 2ef8ca9c-4295-4c58-b682-80c33d7bab6e · outbound

This paper cites Global Convergence in Training Large-Scale Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Global Convergence in Training Large-Scale Transformers

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:46:42.265485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:35.278042Z digest=sha256:a37c2ce6f9c2da184ff1c8c799c3c084d10eeea485b89e363390244b3a43f661

Observation d40ad934-5101-49d1-ae2f-38f0ad56451f · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Better & Faster Large Language Models via Multi-token Prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.454047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.454047Z digest=sha256:9e99a5438da04c6740023d652059085869da687d82f14da21fececaccf9c28fd

Observation 732e7025-32e6-4375-a0f9-6c2d23667c6c · outbound

This paper cites Interleaved group products.SIAM Journal on Computing, 48(2):554–580, 2019.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Interleaved group products.SIAM Journal on Computing, 48(2):554–580, 2019

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.203328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:35.554837Z digest=sha256:5b3452237b209a2daae0575eadd6b08da593789942bbefe1ad78c81dffe9501a

Observation 9fb45c7f-1a84-423b-a32d-f6bf87f6f129 · outbound

This paper cites Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.648813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.648813Z digest=sha256:785ea7e1e92a2fce56ecee1209cfac7eb7016fecfcc693565ddcd7d79fbddf21

Observation 7a1ba3e9-0933-4e24-b9c2-4b1e1948cc63 · outbound

This paper cites In-Context Convergence of Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data In-Context Convergence of Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.756325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.756325Z digest=sha256:e3252cbbf6afac42809868c4b33996d91dc4d330dd4bfe5e12caff7eb01c2fe5

Observation dcb56841-e383-407a-af90-d80388e83cbe · outbound

This paper cites Mas- sively parallel computation: Algorithms and applications.Foundations and Trends® in Optimization, 5(4):340–417, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Mas- sively parallel computation: Algorithms and applications.Foundations and Trends® in Optimization, 5(4):340–417, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.058801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:35.878834Z digest=sha256:ce12b9fa1c7fa1c0401337c5c00423d877643c5d6169c3796dabea5c37aae673

Observation ccca6d86-49a0-4459-8eea-37c03d64ca10 · outbound

This paper cites Vision transformers provably learn spatial structure.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Vision transformers provably learn spatial structure

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.916436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:36.025687Z digest=sha256:a4fe05e9f492415954722d02b5d4a93b63de89e92b80e152cf95fb4550017913

Observation 7d3ab2f4-6822-4c76-afe7-47e5270748c0 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.145807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.145807Z digest=sha256:74df7f19c11fab0e2d9a57eb6bad039f32e8428acd5be8bd93e6818a4262af56

Observation 99bacc77-02fd-418d-9507-034ef2093aaf · outbound

This paper cites Efficient noise-tolerant learning from statistical queries.J.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Efficient noise-tolerant learning from statistical queries.J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.776875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:36.275561Z digest=sha256:c93b7d78fbf7d57d86ab8e5d09bafdf676d3dde447be6d9e0570e129e2120435

Observation 2322fda7-1f28-4a84-bb51-4dec2ba0ffc8 · outbound

This paper cites Transformers Provably Solve Parity Efficiently with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Provably Solve Parity Efficiently with Chain of Thought

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.416475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.416475Z digest=sha256:df43a8bb51ae7ebbd633eaf84dea9396a8681244cf76d59fdccd88e88aed90b8

Observation dd18b5f5-62d5-4173-bc4b-3b016e719ba7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Adam: A Method for Stochastic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.536688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.536688Z digest=sha256:c147a7ee2b92c979920dd3f2cdf6b9431788d3a3fa89644fca2338ed71b07eb0

Observation be79f8cf-1577-4a53-8f21-a77e162349bc · outbound

This paper cites Learning to reason and memorize with self-notes.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning to reason and memorize with self-notes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.652195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:36.648732Z digest=sha256:519ebdbc6be79ca4547222a602de39b6b668739c3aaab0ff62088915142a3bff

Observation 070b5a5f-cc81-49cf-bbce-bb5df2caf729 · outbound

This paper cites Solving quantitative reasoning problems with language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving quantitative reasoning problems with language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.509475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:36.789484Z digest=sha256:9d675a0c47e044f2fd9a9ced7da5a6605b0d82199d6b56c48a1bc7c3c3a12c72

Observation a1257c2b-f6eb-4450-91f6-91492b32782b · outbound

This paper cites How do transformers learn topic structure: Towards a mechanistic understanding.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How do transformers learn topic structure: Towards a mechanistic understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.366060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:36.913612Z digest=sha256:08e57deee6960826e076dc87246fb0320c19a8d245c3108f3740407b031f3665

Observation a5719318-da93-4fef-a358-989d748c5d6b · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.044242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.044242Z digest=sha256:90eb1cb965e728d3e49e9616d690896838344c9a3bedcf2d98c3c78488b80406

Observation 4d5d60aa-d7d8-4822-99a7-c91d5121ced5 · outbound

This paper cites Let's Verify Step by Step.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Let's Verify Step by Step

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.158871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.158871Z digest=sha256:be0662b60d221e6d6d668b5ad1914a237fc734b8dbde59d1fbc8eb3b88c1e3a3

Observation e23e46eb-224b-41a4-8744-80dcfddca948 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning Compositional Functions with Transformers from Easy-to-Hard Data DeepSeek-V3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.256187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.256187Z digest=sha256:1a8be654afcca6ce38be7b35e6a38baa23c8c47cb88ade3005a244f4682b3fd6

Observation d5258745-2be9-4dfb-b9d2-4182af1f2f97 · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.210499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:37.413474Z digest=sha256:4fe201b017a29862fd0001968c56f8e3fd72ed6a3fb146369523793be64eed75

Observation dedc056e-1367-4bac-a297-f4259dc3a96f · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Learning Compositional Functions with Transformers from Easy-to-Hard Data RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.549144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.549144Z digest=sha256:44ce0de8a699dd2564a5141550deaf5c24da0b452010b312024ca63537549a4b

Observation a805c674-e3e4-496c-b700-826fa40ccb80 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The Expressive Power of Transformers with Chain of Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.657775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.657775Z digest=sha256:bbbeb7f3660f01e884d397b2010a51628382869c4e6d284ebae1bbc853c3d21c

Observation ba40d699-5080-47eb-b246-71d1a32e7455 · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.Transactions of the Association for Computational Linguistics, 11:531–545, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The parallelism tradeoff: Limitations of log-precision transformers.Transactions of the Association for Computational Linguistics, 11:531–545, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.789402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.789402Z digest=sha256:5ff35afb4ee5fad04cd5357202bb3e5f601e8b8e7d73c124cb49692570a97706

Observation 5255c7ec-42a6-43ae-b389-7051503a870d · outbound

This paper cites How transformers learn causal structure with gradient descent.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How transformers learn causal structure with gradient descent

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.972428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:37.884202Z digest=sha256:5356d8f82271b64233c3d3a9b5c431d6ba820385ec00a9b0fbb2f5be9b3f83f5

Observation d36466a5-8576-412c-b84a-5474d5d7895d · outbound

This paper cites Understanding Factual Recall in Transformers via Associative Memories.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Factual Recall in Transformers via Associative Memories

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.010378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.010378Z digest=sha256:b1add74052db06b6440d5813a11122c4a398df29e19268f7148073a90c9debba

Observation 367d7a22-b668-4a37-b725-678b2bc061ab · outbound

This paper cites Rounds in communication complexity revisited.SIAM Journal on Computing, 22(1):211–219, 1993.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Rounds in communication complexity revisited.SIAM Journal on Computing, 22(1):211–219, 1993

Reference 45

Resolution
verified exact
doi, observed 2026-08-07T12:46:40.801870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:38.136778Z digest=sha256:7c38d2e9c7ec09a348cab6ef1ef102e187a1375143056b985a99b57594a0a7a1

Observation 45205473-97b1-49c4-b5fb-f8495ed7503f · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.260594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.260594Z digest=sha256:1bcac6d53912a86c6432741a9b068ccef7797c013aee026e5da941b0e382f3ff

Observation 53ea21a2-0bc5-4f93-9ba7-922355824d0e · outbound

This paper cites In-context learning and induction heads.Transformer Cir- cuits Thread, 2022.

Learning Compositional Functions with Transformers from Easy-to-Hard Data In-context learning and induction heads.Transformer Cir- cuits Thread, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.790112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:38.397937Z digest=sha256:6fd24f8983ee6f9b835e4676227d387147ae360d3e3de8c3ef83b1ca722fdf46

Observation ffc1f16e-3425-486b-bab0-270f6c72b8ff · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Progressive distillation induces an implicit curriculum

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.497645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.497645Z digest=sha256:1916e8040253c717d0fcca2917dd3c114da108033d91dffb8048a432ccc98aaf

Observation de5a35ec-1470-4b72-9c7b-c858563207da · outbound

This paper cites Papadimitriou and Michael Sipser.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Papadimitriou and Michael Sipser

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.611842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.611842Z digest=sha256:5b5e3ad51caaa08d6b68d1eb2c5c858c370c0d951fecbe5a6ab42e13bc4db534

Observation e56752ed-55ec-4177-885f-087b1798635a · outbound

This paper cites On limitations of the trans- former architecture.

Learning Compositional Functions with Transformers from Easy-to-Hard Data On limitations of the trans- former architecture

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.565431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:38.755935Z digest=sha256:36515ac8aba7ad3135df9ed34cc4f0722fd1f3cd397364548d50229593fa0d3f

Observation 8f8dcf86-ec50-4dfd-8d3c-c5265be5cb83 · outbound

This paper cites Learning and transferring sparse contextual bi- grams with linear transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and transferring sparse contextual bi- grams with linear transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.433429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:38.862958Z digest=sha256:a7d830fc2dd954c81e4236f354b5c0229eeaef3ae1d98cff72837eac0ddc2cc7

Observation a0aaa9a3-4f95-4cec-b64e-24a101960671 · outbound

This paper cites Representational strengths and limitations of transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Representational strengths and limitations of transformers

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.235854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:38.989395Z digest=sha256:f325b723508cc76b046d160270a5ba1378ab0fa70d1b2dc17c7bd20e38b485a2

Observation f8097cbb-e197-4a3b-b5ea-cb4eb9dceede · outbound

This paper cites Understanding Transformer Reasoning Capabilities via Graph Algorithms.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Transformer Reasoning Capabilities via Graph Algorithms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.144554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.144554Z digest=sha256:bd6498ec6df95e3dc0c8fc842c5d93f9d5789f0e16ee763c51359c564c911c11

Observation e1f0a159-eac9-4ca1-8d46-d41662a12b6f · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers, parallel computation, and logarithmic depth

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.078056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.220249Z digest=sha256:8381b2ce8bf54a1a39e4422ccaeb1e222782c19fb4d0bd17eae6c3b08e976e60

Observation 9792657b-2058-465f-9fe8-172cdcf3327b · outbound

This paper cites End-to-end memory networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data End-to-end memory networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.894174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.313177Z digest=sha256:87fdd05a465479b824f9cf46246ee5eae125f8ceb90c786226c7f9c0ade9ef1b

Observation 19db7bda-130f-4046-83f8-0764f09a3f47 · outbound

This paper cites Characterizing statistical query learning: simplified notions and proofs.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Characterizing statistical query learning: simplified notions and proofs

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.758978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.420028Z digest=sha256:877da5c85209c6795fa163eacf8a0a30032e53687008510f3d76701e992308d2

Observation 36aa5752-892f-40e3-ba6c-081c9d010c5c · outbound

This paper cites Scan and snap: Understanding training dynamics and token composition in 1-layer transformer.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Scan and snap: Understanding training dynamics and token composition in 1-layer transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.593968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.519276Z digest=sha256:a7d0998943cf7e605b24e6e6ebb01c91381ddcce077537bc53495544c93042a8

Observation 22e35c1f-1df0-4579-a963-34b49678ef39 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving math word problems with process- and outcome-based feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.598022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.598022Z digest=sha256:80448420ebf935bbce46189d2ea553fba136ae0d930a604fb49bf6ea3f1e9efb

Observation 2259d854-e282-4a79-aa47-a6c1ad55e9a9 · outbound

This paper cites an unresolved cited work.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:46:44.399090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.698128Z digest=sha256:f7d14ea19a71ecb3989f30317fe94fe1920ac1d89e3c5b23579cec66e113d7e0

Observation 0da4c9a8-52a1-46f5-aef6-94ee1512ebdf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain-of-thought prompting elicits reasoning in large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.220213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:39.771667Z digest=sha256:22339546f906b193a03b616e58869f0a515eb18a97f02cac1c0c837271a8ce75

Observation 4c59a0ee-c65c-4fe3-8215-7a17f3099a43 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Learning Compositional Functions with Transformers from Easy-to-Hard Data From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.840714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.840714Z digest=sha256:255f3ff9b714d067269fd65b640b373f01c065f8aaa1350f3239a37e1d098487

Observation b7d6e0cb-cbc7-4f78-99c0-39d9167d5ef4 · outbound

This paper cites Memory Networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Memory Networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.905563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.905563Z digest=sha256:ee5224c8fb801fb10e5e6d4d04b264e549ebfe026886bc54d2bb1066c5b0eb1f

Observation 02cb2c74-8a26-46ff-a8aa-d8384dd7a2a4 · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.984911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.984911Z digest=sha256:8462c35aee37c5963010646a9282af5dd7e877f6b3fd7a841d513dc49030b5de

Observation d0bcd8b8-c830-42e1-9a0c-ce3c25732862 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.044810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.056872Z digest=sha256:555cd020c0051d6fde115f0be2ff56f510871149c480eabce3d641ea832d0863

Observation 2f6f0c7b-8d67-432a-b51b-fb187d7f4894 · outbound

This paper cites Do Large Language Models Latently Perform Multi-Hop Reasoning?.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Do Large Language Models Latently Perform Multi-Hop Reasoning?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:40.093818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:40.093818Z digest=sha256:6c2e4e56f84b6f5cf28943d8a3f53b505cdd154a13bd6abb027b53ed103b5b9c

Observation 6a3e5602-bead-4bd9-99ca-23dab1739c35 · outbound

This paper cites Tree of thoughts: deliberate problem solving with large language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Tree of thoughts: deliberate problem solving with large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.858525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.210788Z digest=sha256:7adf2e0fc31e80c68b0a3920133e055c603ee6d7f18d15cb71ea47a5c319facf

Observation 3c0f99ec-21a5-482c-b5d4-71d77c1f7c13 · outbound

This paper cites Pointer chasing via triangular discrimination.Combinatorics, Probability and Computing, 29(4):485–494, 2020.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Pointer chasing via triangular discrimination.Combinatorics, Probability and Computing, 29(4):485–494, 2020

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.665821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.299772Z digest=sha256:7b21a119705e5f5b89bc56db7b6a3301c80b9eb2748ad6ae9fcb05d620ed58e6

Observation b2fc7f14-4432-4a2a-a85f-c240aa9b09e9 · outbound

This paper cites Trained Transformers Learn Linear Models In-Context.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Trained Transformers Learn Linear Models In-Context

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:40.385953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:40.385953Z digest=sha256:8e0bcff42109d1b665fbbd450a4e4d9d17ee74a02c9709c404f676652bf6c248

Observation 48c3ed10-53dd-4f8b-98cf-3c616de2aa69 · outbound

This paper cites an unresolved cited work.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:46:43.471531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.475130Z digest=sha256:e3bf0ce6d4d46420c9547c2a6449c385aec71a05c4ea90032a13ab6e141be94f

Observation 6efa7547-edbe-489e-869f-cbc4f2d7184a · outbound

This paper cites Since the softmax in layer ℓ is nearly saturated and thus close to one-hot, the update norm is very small and can be bounded as noise terms.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Since the softmax in layer ℓ is nearly saturated and thus close to one-hot, the update norm is very small and can be bounded as noise terms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.368846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.545158Z digest=sha256:8ad01d8b2d1c45a6f6a0c4bb28bea8996f766ca9b8ce2e8ffd32d4c3fdb5d1c2

Observation 192e0eaa-499b-4984-bd8d-e32f3cb90dc2 · outbound

This paper cites That means in the tth gradient step, there will be some small gradient updates for later layersℓ′ ≥t introduced by the perturbation.

Learning Compositional Functions with Transformers from Easy-to-Hard Data That means in the tth gradient step, there will be some small gradient updates for later layersℓ′ ≥t introduced by the perturbation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.194198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:46:40.614748Z digest=sha256:d70b15699f5aed6f7053aecaf5dd3e809f0fefa13dd40e250a2d9d032347a09d

Pith citing papers

Observation 1e43c969-1d4f-439d-8343-a97345183408 · inbound

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence cites this paper.

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:15.603859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:15.603859Z digest=sha256:81d7d462d270125ab93029f2e2547c8fa44905f46ac514bc43466f0135729204

Observation 67828437-2bbd-46e5-b46c-cd501bce0c1b · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:43:11.763654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:f1b4b8ddb55966cd350da11c29474ea6ce8efefaed22b39aa768dcdb20204756

Observation feed43f0-78ec-4421-a44f-ce39f10e4457 · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:49.270195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:49.270195Z digest=sha256:00075c022eb04e174cdec2b7fcb7f975c0bd663820fb63c53ab8296c555a0e2d

Observation 7ff28ce0-6fcb-4a8c-accc-578779f06471 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 186

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.310219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:970e1991ccdeaafb94885b3861cc44d76c4a39d58a452391d413213781974e5b

Observation 0b1e143f-b234-4151-8cd5-0c79778c1c00 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:12.841539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:12.841539Z digest=sha256:3d3366439cd050c01ba63e9d77151b087e89c07821bb4870822f95efb8cb4f10

Observation 3340179c-4d68-4d1f-8b84-d172d3085ec2 · inbound

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge cites this paper.

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:44.481599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:44.481599Z digest=sha256:2310f31cb797f1a55dd28d4b50c485f789b1996030e393050af668939b921d08

Observation c1fafadd-a6b4-4052-9480-75040c60f282 · inbound

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cites this paper.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:38:16.390943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T23:37:33.106390Z digest=sha256:5df8ed16e51297b9984b45a69d606a6122677c4d0bf327cf8cac06ac1d41c22e

Observation 2e0e0277-07c5-4180-b57c-d9ec1f0787ec · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.436718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:2316a172e7d5b873fabd6ad665d34e26e457358c172ef6b6d6d433b6cf0fade4

Observation 26962014-0ff3-463c-a610-90aef65b491b · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:5e07507a9d2012b1662671e2669b5853259769ad527762ca08d80c6c0a66ba72

Observation 7a4a587c-aaa8-4cc0-8415-c2fd22c76ffe · inbound

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent cites this paper.

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:09.597626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T12:22:27.265973Z digest=sha256:24a104749fc668ceb1d3f4890fc31a72507e3d62e595a63a892ec1adb18f082f

Observation 3836b10a-2e80-4a5b-9cc5-52054adc3b64 · inbound

Transformers Provably Learn to Internalize Chain-of-Thought cites this paper.

Transformers Provably Learn to Internalize Chain-of-Thought Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:33:30.600195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T14:29:10.010212Z digest=sha256:0eebf4d81f89eb20f547dc2920d5cfa2598991073af6775c2532daf90c8d321f