Pith. sign in

Paper Citation Record · LEDGER

Learning Compositional Functions with Transformers from Easy-to-Hard Data

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 11 inbound Pith citation observations for arXiv:2505.23683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23683 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:40.614748Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:15.603859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:33:30.598717Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c791e2bc-b13d-47aa-97a8-afcc6650bf6b · outbound

This paper cites The staircase property: How hierarchical structure can guide deep learning.Advances in Neural Information Processing Systems, 34:26989–27002, 2021.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The staircase property: How hierarchical structure can guide deep learning.Advances in Neural Information Processing Systems, 34:26989–27002, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:32.494326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:32.494326Z digest=sha256:e302b1de3082391b3123ce0402d62575e57caa77895ba5b1e0d838571620a051

Observation ddd8ead5-d8c7-406c-80cd-4067f3f7b8da · outbound

This paper cites The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.442515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:32.596789Z digest=sha256:12754cf03036c8acbaaa4d9c27f9d88fb7dc4f918584b89558e4ac1bc30ac5a7

Observation a03d8f2f-fec3-480a-9a4c-524dcb10114c · outbound

This paper cites Provable advantage of curriculum learning on parity targets with mixed inputs.Advances in Neural Information Processing Systems, 36:24291–24321, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Provable advantage of curriculum learning on parity targets with mixed inputs.Advances in Neural Information Processing Systems, 36:24291–24321, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:32.708882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:32.708882Z digest=sha256:d2766ba4ecb4d4939b9d177e17a3d89afd90933cbac7d277c9b40d96c416d9dd

Observation 831013f1-6797-4024-ab3a-b34907bde771 · outbound

This paper cites How far can transformers reason? The locality barrier and inductive scratchpad.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How far can transformers reason? The locality barrier and inductive scratchpad

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.275011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:32.865237Z digest=sha256:6e26954c2c2fcdb1d6ccd50c561ddfe41eff18bb53880bd39581001d3ceea9f6

Observation 773f1920-9aa9-4517-95ed-c7f90f9bde19 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.160610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:32.939839Z digest=sha256:b9287a33ad1e5d627c11a8466295b6eca7dd89922eedfc18b61400fd16df142e

Observation c86c7eff-538c-4a24-8d82-21a2ad229fac · outbound

This paper cites Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:48.033495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:33.040341Z digest=sha256:98ef5ca6cd35ce95972ca10f3609a542ba6fbbf6f859d8b675a74f15b21cf067

Observation 4ab05c7b-4394-466c-b552-cbf39e6178de · outbound

This paper cites The pitfalls of next-token prediction.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The pitfalls of next-token prediction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.870785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:33.173794Z digest=sha256:9743f43895bd3b7d75e609cec2327eb146647d9912e8c5d95c26c596337110b9

Observation b6ed973d-856d-422d-89bd-59af7f89cea5 · outbound

This paper cites Curriculum learning.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Curriculum learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.334050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.334050Z digest=sha256:2f104dc1243020dd87379e0d6e2f96c43ffd98378d86c76bd2d2b06235b2a9f0

Observation 36d3242c-1566-4cde-b737-5d22627eb4f5 · outbound

This paper cites Separations in the Representational Capabilities of Transformers and Recurrent Architectures.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Separations in the Representational Capabilities of Transformers and Recurrent Architectures

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:46:42.789421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:33.485218Z digest=sha256:5a91d782e04db9dc86446d91acfc9035bb5f995863ca2320efd51442d2578de2

Observation ccb8cf36-5b8b-4f33-a4be-6228c304285c · outbound

This paper cites Birth of a transformer: A memory viewpoint.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Birth of a transformer: A memory viewpoint

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.760483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:33.649966Z digest=sha256:397c82f3c5a47585ff0b0457e4a9aa3058e911de99b4691947f19ee4d7e74dc4

Observation c834b064-facb-4bf2-9fb7-c7c5926c5ae1 · outbound

This paper cites Data distributional properties drive emergent in-context learning in transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Data distributional properties drive emergent in-context learning in transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.608752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:33.769852Z digest=sha256:3d2bf7bc75d38faa257a0778a73190e8c68304be2c5bb2ebfb0dc8a41f8a3710

Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · outbound

This paper cites Theoretical limitations of multi-layer Transformer.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.978754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.978754Z digest=sha256:204e6b826ea53e67fb2ea10b95b54b720a40e7d40f9624c3c7d0485ea2a7fe7c

Observation a597f3a7-8913-43cb-bfa4-167279020fab · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.110434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.110434Z digest=sha256:26331ba4018d0d3df3b478b197990f0cae8c7a6969efffdd1994e15fedaf107a

Observation b134bc0c-7e37-4fb7-824d-df2c5462488e · outbound

This paper cites Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.250512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.250512Z digest=sha256:2b6568d9cf93fb71a610eb51a19d04ddfc418de380b3d51427fc28cb227b11fa

Observation 3ed040cd-8986-4385-b60e-f88d2bf4c500 · outbound

This paper cites Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.394535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.394535Z digest=sha256:522c942e40202da936577e0e850f2f544c685033dbf71c877a7ba18a84e5748b

Observation 8bde2b42-80a5-4895-a919-a11e66fefed9 · outbound

This paper cites Neural networks can learn represen- tations with gradient descent.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Neural networks can learn represen- tations with gradient descent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.489754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.489754Z digest=sha256:a9fb64f7624d3994bf4d973e67d73be3920824b2e1dc85fbcead6576b37eb5f2

Observation b5054e96-626f-42fc-be1a-26f48263814f · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Learning Compositional Functions with Transformers from Easy-to-Hard Data From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.662611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.662611Z digest=sha256:0fb737d4e2093e6ac19a598e24c722d259c30ebdfe76a91b22e7664ab7651f0b

Observation 9bfa3d55-b086-4846-a561-d6cee1fbf8ce · outbound

This paper cites The Llama 3 Herd of Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.805557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.805557Z digest=sha256:b11c5e6d948da7d177044c3bca9c6359cd6f44f4082515eef185d16af5f808db

Observation c841ccb3-8b76-48cc-a404-53bbb7898b27 · outbound

This paper cites Faith and Fate: Limits of Transformers on Compositionality.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Faith and Fate: Limits of Transformers on Compositionality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.917688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.917688Z digest=sha256:04e57592a22946ae5867e85dfb5e28179ca8e14986d512cc6d53a283406ca2d7

Observation e4bc1d60-d879-4948-895a-d503ea825cc7 · outbound

This paper cites Learning and development in neural networks: The importance of starting small.Cognition, 48(1):71–99, 1993.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and development in neural networks: The importance of starting small.Cognition, 48(1):71–99, 1993

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.061289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.061289Z digest=sha256:628be2d7659f2b4f7d9a973dcbf879b5743ddf50cdb0771f3c5b293cbc203f15

Observation b543bd7a-7bda-4750-b189-27547524ec73 · outbound

This paper cites Towards revealing the mystery behind chain of thought: a theoretical perspective.Advances in Neural Information Processing Systems, 36, 2024.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards revealing the mystery behind chain of thought: a theoretical perspective.Advances in Neural Information Processing Systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.382613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:35.155644Z digest=sha256:9ee2758c7324d2d4a33ef24735ebdcd400d86b7d19720d6d61ff8eec6e8c035b

Observation 2ef8ca9c-4295-4c58-b682-80c33d7bab6e · outbound

This paper cites Global Convergence in Training Large-Scale Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Global Convergence in Training Large-Scale Transformers

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:46:42.265485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:35.278042Z digest=sha256:93f2fc0e17060009c8343232091c4b72d1a3d78f299419a2ba39806bbc8e66a8

Observation d40ad934-5101-49d1-ae2f-38f0ad56451f · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Better & Faster Large Language Models via Multi-token Prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.454047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.454047Z digest=sha256:9e99a5438da04c6740023d652059085869da687d82f14da21fececaccf9c28fd

Observation 732e7025-32e6-4375-a0f9-6c2d23667c6c · outbound

This paper cites Interleaved group products.SIAM Journal on Computing, 48(2):554–580, 2019.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Interleaved group products.SIAM Journal on Computing, 48(2):554–580, 2019

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.203328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:35.554837Z digest=sha256:104637e7e714357ed3bcc8267c26659e977d99eed51518652f6e112a9a008042

Observation 9fb45c7f-1a84-423b-a32d-f6bf87f6f129 · outbound

This paper cites Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.648813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.648813Z digest=sha256:955f9f322dd3aacb1d0e10a5ab7f067aa701e43544332b6689e88af11b0d94dd

Observation 7a1ba3e9-0933-4e24-b9c2-4b1e1948cc63 · outbound

This paper cites In-Context Convergence of Transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data In-Context Convergence of Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:35.756325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:35.756325Z digest=sha256:e3252cbbf6afac42809868c4b33996d91dc4d330dd4bfe5e12caff7eb01c2fe5

Observation dcb56841-e383-407a-af90-d80388e83cbe · outbound

This paper cites Mas- sively parallel computation: Algorithms and applications.Foundations and Trends® in Optimization, 5(4):340–417, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Mas- sively parallel computation: Algorithms and applications.Foundations and Trends® in Optimization, 5(4):340–417, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:47.058801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:35.878834Z digest=sha256:7fb79a0c5ee221371fbab8b43f01b1ef41754d59d71d218438e35fd1d6eedf58

Observation ccca6d86-49a0-4459-8eea-37c03d64ca10 · outbound

This paper cites Vision transformers provably learn spatial structure.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Vision transformers provably learn spatial structure

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.916436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:36.025687Z digest=sha256:5bdc66d99b6c657faea703d0e970ed12af32bf8b358701ac070e4dcfcaf31c12

Observation 7d3ab2f4-6822-4c76-afe7-47e5270748c0 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.145807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.145807Z digest=sha256:74df7f19c11fab0e2d9a57eb6bad039f32e8428acd5be8bd93e6818a4262af56

Observation 99bacc77-02fd-418d-9507-034ef2093aaf · outbound

This paper cites Efficient noise-tolerant learning from statistical queries.J.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Efficient noise-tolerant learning from statistical queries.J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.776875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:36.275561Z digest=sha256:1a65db3a3e3e2b3deb67ea8a7c5c35c21c71643a7fb862ab7d6c18ef918a579b

Observation 2322fda7-1f28-4a84-bb51-4dec2ba0ffc8 · outbound

This paper cites Transformers Provably Solve Parity Efficiently with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Provably Solve Parity Efficiently with Chain of Thought

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.416475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.416475Z digest=sha256:0aaff65347051005e9f89e9c104d7c4d4aa207e926a85298ad8d08e3b09cd0ba

Observation dd18b5f5-62d5-4173-bc4b-3b016e719ba7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Adam: A Method for Stochastic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.536688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.536688Z digest=sha256:739d5c274cf473c8ce07a55415d502c6002a3c9cc847f912828d94ca7598c81b

Observation be79f8cf-1577-4a53-8f21-a77e162349bc · outbound

This paper cites Learning to reason and memorize with self-notes.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning to reason and memorize with self-notes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.652195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:36.648732Z digest=sha256:00197035b52792c0520f03448ff5547bfe50e8ebc0574e670f759b250340a69e

Observation 070b5a5f-cc81-49cf-bbce-bb5df2caf729 · outbound

This paper cites Solving quantitative reasoning problems with language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving quantitative reasoning problems with language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.509475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:36.789484Z digest=sha256:449e3f1f114f9e5708305c60cc65ab91b53278117e3fead56abe35a020a29936

Observation a1257c2b-f6eb-4450-91f6-91492b32782b · outbound

This paper cites How do transformers learn topic structure: Towards a mechanistic understanding.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How do transformers learn topic structure: Towards a mechanistic understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.366060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:36.913612Z digest=sha256:a822d5dab7320978e5a92f96e6a3e068fe875f313e4ceb41e265f7447abe9a39

Observation a5719318-da93-4fef-a358-989d748c5d6b · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.044242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.044242Z digest=sha256:90eb1cb965e728d3e49e9616d690896838344c9a3bedcf2d98c3c78488b80406

Observation 4d5d60aa-d7d8-4822-99a7-c91d5121ced5 · outbound

This paper cites Let's Verify Step by Step.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Let's Verify Step by Step

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.158871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.158871Z digest=sha256:be0662b60d221e6d6d668b5ad1914a237fc734b8dbde59d1fbc8eb3b88c1e3a3

Observation e23e46eb-224b-41a4-8744-80dcfddca948 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning Compositional Functions with Transformers from Easy-to-Hard Data DeepSeek-V3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.256187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.256187Z digest=sha256:1a8be654afcca6ce38be7b35e6a38baa23c8c47cb88ade3005a244f4682b3fd6

Observation d5258745-2be9-4dfb-b9d2-4182af1f2f97 · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:46.210499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:37.413474Z digest=sha256:463e73fc7cf1e4861b79cdb7dee71932099200ecf21367e118a66dea7a69b48f

Observation dedc056e-1367-4bac-a297-f4259dc3a96f · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Learning Compositional Functions with Transformers from Easy-to-Hard Data RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.549144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.549144Z digest=sha256:44ce0de8a699dd2564a5141550deaf5c24da0b452010b312024ca63537549a4b

Observation a805c674-e3e4-496c-b700-826fa40ccb80 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The Expressive Power of Transformers with Chain of Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.657775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.657775Z digest=sha256:bbbeb7f3660f01e884d397b2010a51628382869c4e6d284ebae1bbc853c3d21c

Observation ba40d699-5080-47eb-b246-71d1a32e7455 · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.Transactions of the Association for Computational Linguistics, 11:531–545, 2023.

Learning Compositional Functions with Transformers from Easy-to-Hard Data The parallelism tradeoff: Limitations of log-precision transformers.Transactions of the Association for Computational Linguistics, 11:531–545, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:37.789402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:37.789402Z digest=sha256:5ff35afb4ee5fad04cd5357202bb3e5f601e8b8e7d73c124cb49692570a97706

Observation 5255c7ec-42a6-43ae-b389-7051503a870d · outbound

This paper cites How transformers learn causal structure with gradient descent.

Learning Compositional Functions with Transformers from Easy-to-Hard Data How transformers learn causal structure with gradient descent

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.972428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:37.884202Z digest=sha256:46dfbd34cddd6330b26d79e088322be104ff2978971bc79e442d444d5737712f

Observation d36466a5-8576-412c-b84a-5474d5d7895d · outbound

This paper cites Understanding Factual Recall in Transformers via Associative Memories.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Factual Recall in Transformers via Associative Memories

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.010378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.010378Z digest=sha256:f1592a735382da8291d280ae12b73864de03f09b440fd35addd5dc0a76d40553

Observation 367d7a22-b668-4a37-b725-678b2bc061ab · outbound

This paper cites Rounds in communication complexity revisited.SIAM Journal on Computing, 22(1):211–219, 1993.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Rounds in communication complexity revisited.SIAM Journal on Computing, 22(1):211–219, 1993

Reference 45

Resolution
verified exact
doi, observed 2026-08-07T12:46:40.801870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:38.136778Z digest=sha256:8f0289138a9fa4a26f3c49edf67792f4d6dde67ae6fa1943a427a36bc5cfeb99

Observation 45205473-97b1-49c4-b5fb-f8495ed7503f · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.260594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.260594Z digest=sha256:1bcac6d53912a86c6432741a9b068ccef7797c013aee026e5da941b0e382f3ff

Observation 53ea21a2-0bc5-4f93-9ba7-922355824d0e · outbound

This paper cites In-context learning and induction heads.Transformer Cir- cuits Thread, 2022.

Learning Compositional Functions with Transformers from Easy-to-Hard Data In-context learning and induction heads.Transformer Cir- cuits Thread, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.790112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:38.397937Z digest=sha256:991e9557b1a69e5aaba2d75857bd44bfa09fb3d1423a0a6c41c69bee9f423d80

Observation ffc1f16e-3425-486b-bab0-270f6c72b8ff · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Progressive distillation induces an implicit curriculum

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.497645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.497645Z digest=sha256:1916e8040253c717d0fcca2917dd3c114da108033d91dffb8048a432ccc98aaf

Observation de5a35ec-1470-4b72-9c7b-c858563207da · outbound

This paper cites Papadimitriou and Michael Sipser.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Papadimitriou and Michael Sipser

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:38.611842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:38.611842Z digest=sha256:5b5e3ad51caaa08d6b68d1eb2c5c858c370c0d951fecbe5a6ab42e13bc4db534

Observation e56752ed-55ec-4177-885f-087b1798635a · outbound

This paper cites On limitations of the trans- former architecture.

Learning Compositional Functions with Transformers from Easy-to-Hard Data On limitations of the trans- former architecture

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.565431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:38.755935Z digest=sha256:85e9e3adcddc47ec9e576361b5b2012bcec15644aa5f6995a5337444c760f0e5

Observation 8f8dcf86-ec50-4dfd-8d3c-c5265be5cb83 · outbound

This paper cites Learning and transferring sparse contextual bi- grams with linear transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and transferring sparse contextual bi- grams with linear transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.433429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:38.862958Z digest=sha256:eb7cffc23810f6675390e3177ce70b5df6c17f887b5b0d1eaf4168cd5d0cab61

Observation a0aaa9a3-4f95-4cec-b64e-24a101960671 · outbound

This paper cites Representational strengths and limitations of transformers.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Representational strengths and limitations of transformers

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.235854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:38.989395Z digest=sha256:a22adf0bc13bce0862806dad16e201419d8d9c04419f70a6710f4ee8670ad761

Observation f8097cbb-e197-4a3b-b5ea-cb4eb9dceede · outbound

This paper cites Understanding Transformer Reasoning Capabilities via Graph Algorithms.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Transformer Reasoning Capabilities via Graph Algorithms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.144554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.144554Z digest=sha256:bd6498ec6df95e3dc0c8fc842c5d93f9d5789f0e16ee763c51359c564c911c11

Observation e1f0a159-eac9-4ca1-8d46-d41662a12b6f · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers, parallel computation, and logarithmic depth

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:45.078056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.220249Z digest=sha256:665462873cbf129e8f731845d30d17f4d1932e90449aa56606dc1bd85808d56f

Observation 9792657b-2058-465f-9fe8-172cdcf3327b · outbound

This paper cites End-to-end memory networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data End-to-end memory networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.894174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.313177Z digest=sha256:5009d7306ccafb94a9ec6951c53abfd78dee485f205c6bf970e1b1e2060effff

Observation 19db7bda-130f-4046-83f8-0764f09a3f47 · outbound

This paper cites Characterizing statistical query learning: simplified notions and proofs.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Characterizing statistical query learning: simplified notions and proofs

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.758978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.420028Z digest=sha256:8511254d7171d996e129216b12c6224779286cb8275aa37533fc140018eebddc

Observation 36aa5752-892f-40e3-ba6c-081c9d010c5c · outbound

This paper cites Scan and snap: Understanding training dynamics and token composition in 1-layer transformer.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Scan and snap: Understanding training dynamics and token composition in 1-layer transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.593968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.519276Z digest=sha256:4ec491559d4d7aaaca08c8bf2f3d2ec522ef63ef4fa195ce3f71e5c2360f1470

Observation 22e35c1f-1df0-4579-a963-34b49678ef39 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving math word problems with process- and outcome-based feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.598022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.598022Z digest=sha256:80448420ebf935bbce46189d2ea553fba136ae0d930a604fb49bf6ea3f1e9efb

Observation 2259d854-e282-4a79-aa47-a6c1ad55e9a9 · outbound

This paper cites an unresolved cited work.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:46:44.399090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.698128Z digest=sha256:5f60c88395f1b4ec627e0afb82ca2c6c006e93e079830943bba5d2f0d0b633dc

Observation 0da4c9a8-52a1-46f5-aef6-94ee1512ebdf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain-of-thought prompting elicits reasoning in large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.220213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:39.771667Z digest=sha256:74c9a21a7fe5f61f0c6f7c6f46f26ddd6a41648c4ccde193681b94124ad5abec

Observation 4c59a0ee-c65c-4fe3-8215-7a17f3099a43 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Learning Compositional Functions with Transformers from Easy-to-Hard Data From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.840714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.840714Z digest=sha256:255f3ff9b714d067269fd65b640b373f01c065f8aaa1350f3239a37e1d098487

Observation b7d6e0cb-cbc7-4f78-99c0-39d9167d5ef4 · outbound

This paper cites Memory Networks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Memory Networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.905563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.905563Z digest=sha256:ee5224c8fb801fb10e5e6d4d04b264e549ebfe026886bc54d2bb1066c5b0eb1f

Observation 02cb2c74-8a26-46ff-a8aa-d8384dd7a2a4 · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:39.984911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:39.984911Z digest=sha256:8462c35aee37c5963010646a9282af5dd7e877f6b3fd7a841d513dc49030b5de

Observation d0bcd8b8-c830-42e1-9a0c-ce3c25732862 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:44.044810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.056872Z digest=sha256:eb31135f793c3f486c492282fa426ef8bd726a3440ab9842c1a8f1910adfcfc4

Observation 2f6f0c7b-8d67-432a-b51b-fb187d7f4894 · outbound

This paper cites Do Large Language Models Latently Perform Multi-Hop Reasoning?.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Do Large Language Models Latently Perform Multi-Hop Reasoning?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:40.093818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:40.093818Z digest=sha256:6c2e4e56f84b6f5cf28943d8a3f53b505cdd154a13bd6abb027b53ed103b5b9c

Observation 6a3e5602-bead-4bd9-99ca-23dab1739c35 · outbound

This paper cites Tree of thoughts: deliberate problem solving with large language models.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Tree of thoughts: deliberate problem solving with large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.858525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.210788Z digest=sha256:73092714a1ea9fddb02e17f321c1a7458a4da9e9b80d8eb389d4eb9c18cbaa9c

Observation 3c0f99ec-21a5-482c-b5d4-71d77c1f7c13 · outbound

This paper cites Pointer chasing via triangular discrimination.Combinatorics, Probability and Computing, 29(4):485–494, 2020.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Pointer chasing via triangular discrimination.Combinatorics, Probability and Computing, 29(4):485–494, 2020

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.665821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.299772Z digest=sha256:4002b15daf3fbdca0f771c5a2f9c604504cc1a1298b59954abfbadac7d7be47c

Observation b2fc7f14-4432-4a2a-a85f-c240aa9b09e9 · outbound

This paper cites Trained Transformers Learn Linear Models In-Context.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Trained Transformers Learn Linear Models In-Context

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:40.385953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:40.385953Z digest=sha256:8e0bcff42109d1b665fbbd450a4e4d9d17ee74a02c9709c404f676652bf6c248

Observation 48c3ed10-53dd-4f8b-98cf-3c616de2aa69 · outbound

This paper cites an unresolved cited work.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:46:43.471531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.475130Z digest=sha256:028aeca14eaca4734a4910ceccae8f3de178d9a387908a642bb987864b3d31a2

Observation 6efa7547-edbe-489e-869f-cbc4f2d7184a · outbound

This paper cites Since the softmax in layer ℓ is nearly saturated and thus close to one-hot, the update norm is very small and can be bounded as noise terms.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Since the softmax in layer ℓ is nearly saturated and thus close to one-hot, the update norm is very small and can be bounded as noise terms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.368846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.545158Z digest=sha256:d30aaeb51dad2769e00801b04f694ad6f1aa678be0d54aab7a3394256933f6d2

Observation 192e0eaa-499b-4984-bd8d-e32f3cb90dc2 · outbound

This paper cites That means in the tth gradient step, there will be some small gradient updates for later layersℓ′ ≥t introduced by the perturbation.

Learning Compositional Functions with Transformers from Easy-to-Hard Data That means in the tth gradient step, there will be some small gradient updates for later layersℓ′ ≥t introduced by the perturbation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:43.194198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:46:40.614748Z digest=sha256:837430269fc6960b778568d692abf805aa5798f2f6d9fae1e31f63a1b90e006c

Pith citing papers

Observation 1e43c969-1d4f-439d-8343-a97345183408 · inbound

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence cites this paper.

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:15.603859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:15.603859Z digest=sha256:ae5f2343c67b839a97649974404800fbcc8444ac5f581d472d8bd2865e7f0260

Observation 67828437-2bbd-46e5-b46c-cd501bce0c1b · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:43:11.763654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:0f46f25f95ba9030d40b35da30e8bb12d3ba8bd482f306f4ef703ebb69b8df2e

Observation feed43f0-78ec-4421-a44f-ce39f10e4457 · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:49.270195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:49.270195Z digest=sha256:4980ab655538db8f7fd7e8dd100173f41474be7f33916e335ce2b2fd01217cd8

Observation 7ff28ce0-6fcb-4a8c-accc-578779f06471 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 186

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.310219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:2c522f993b46cf9db7b720856e5d56d86d848ce45e5d54a33fb18ca0dc116383

Observation 0b1e143f-b234-4151-8cd5-0c79778c1c00 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:12.841539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:12.841539Z digest=sha256:3d3366439cd050c01ba63e9d77151b087e89c07821bb4870822f95efb8cb4f10

Observation 3340179c-4d68-4d1f-8b84-d172d3085ec2 · inbound

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge cites this paper.

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:44.481599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:44.481599Z digest=sha256:2310f31cb797f1a55dd28d4b50c485f789b1996030e393050af668939b921d08

Observation c1fafadd-a6b4-4052-9480-75040c60f282 · inbound

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cites this paper.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:38:16.390943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T23:37:33.106390Z digest=sha256:d09f33f4480a80d84346be5f8f59d816293dbd5ff5d0349e29121a95592883e3

Observation 2e0e0277-07c5-4180-b57c-d9ec1f0787ec · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.436718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:47a7b1e29829807b890b19653ab2283328e6f53fe53fe8a879972100942da86a

Observation 26962014-0ff3-463c-a610-90aef65b491b · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:5e07507a9d2012b1662671e2669b5853259769ad527762ca08d80c6c0a66ba72

Observation 7a4a587c-aaa8-4cc0-8415-c2fd22c76ffe · inbound

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent cites this paper.

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:09.597626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T12:22:27.265973Z digest=sha256:0a7cec78a28b25939a9b2451212c7a68d2e1c580d9ee99fa42ffdb5b39ed895e

Observation 3836b10a-2e80-4a5b-9cc5-52054adc3b64 · inbound

Transformers Provably Learn to Internalize Chain-of-Thought cites this paper.

Transformers Provably Learn to Internalize Chain-of-Thought Learning Compositional Functions with Transformers from Easy-to-Hard Data

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:33:30.600195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-29T14:29:10.010212Z digest=sha256:326364ebd899f0fbb09bba07dce400eefe4538a33ef3f19c969b63bc3a6d80c9