Pith. sign in

Paper Citation Record · LEDGER

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

As of 10 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 2 inbound Pith citation observations for arXiv:2502.01694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01694 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:40.815997Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:36:23.845366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:36:23.932030Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50c00ae1-3b43-4ffa-b4e7-ba111481a8b4 · outbound

This paper cites How far can transformers reason? T he locality barrier and inductive scratchpad.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How far can transformers reason? T he locality barrier and inductive scratchpad

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.536637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.569761Z digest=sha256:cb77a4670333cf80bfe74c9e1de5aa7357e0f94eeb6790e2366d9ed220e717d2

Observation e438d207-4bb6-4b9f-896e-5d1448966b19 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.574169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.574169Z digest=sha256:bdc64221ea498487f0a4b0b70f2f6558cbf9aa8290fa42542b26129a3f0d29f2

Observation 24a5fb29-86a9-4465-b1a8-8593739dab93 · outbound

This paper cites Beltr\' a n and C.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Beltr\' a n and C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.528255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.578030Z digest=sha256:5e29f87ae6f237f5c51f1d95ac48a1798364632fdb4fec8d49c0a70ef531bbf0

Observation ac51a1ab-8cde-4ef2-b566-066b90401c98 · outbound

This paper cites Graph of thoughts: solving elaborate problems with large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Graph of thoughts: solving elaborate problems with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.520267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.581505Z digest=sha256:96df1f0a98b19191b01694e41eed02f6090e97d6f8f16e1b1c997f88cfc4d1f6

Observation 0cb9b975-6509-42de-b57d-622616770bd3 · outbound

This paper cites Multi-scale metastable dynamics and the asymptotic stationary distribution of perturbed M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Multi-scale metastable dynamics and the asymptotic stationary distribution of perturbed M arkov chains

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.511374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.585137Z digest=sha256:7efe5a6058fe45cd79360fcea8c0501a7adb8f7a4be4779eab9eadd7d01fca5c

Observation 2b994676-eff7-4a6f-8c49-7003c4e6a366 · outbound

This paper cites Understanding in-context learning in transformers and LLM s by learning to learn discrete functions.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Understanding in-context learning in transformers and LLM s by learning to learn discrete functions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.588486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.588486Z digest=sha256:87f3aa26d39688dcaf2d435a7d1b1d7f2bc35ec94e775fdc088ed417901de245

Observation 6c81d042-285a-44ee-b136-05f8e29d1f98 · outbound

This paper cites Metastable states, quasi-stationary distributions and soft measures.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastable states, quasi-stationary distributions and soft measures

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.495805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.592115Z digest=sha256:6a1b6adbcdba6ad3e337315d796e460e860f456533416f02f9b5c5c49e79a23b

Observation 89100101-25d3-4c9c-86e4-ec4e3829ef17 · outbound

This paper cites Metastability and low lying spectra in reversible M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability and low lying spectra in reversible M arkov chains

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.486093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.595430Z digest=sha256:3e7e4b63fd582b8e2d62ab6a5c1f94077bd3700f10ae5100e106ee1de59d74a2

Observation 9dd31d78-eb4b-4ea7-b163-a0b46229129f · outbound

This paper cites On using extended statistical queries to avoid membership queries.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation On using extended statistical queries to avoid membership queries

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.476089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.598732Z digest=sha256:78f91a0d8975ecf9e95d5bce31c09a8965424cac283f3935b3f274fc689d9f50

Observation 54e4cb9c-3218-473e-a04b-8ab501f75d64 · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.466392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.602193Z digest=sha256:857ed63d54f919e559a4a72169fa7ebd457f57c77ffccfcce29d50cb03c4a4ea

Observation 19df1bd3-babe-439a-aa57-3fb4c06146de · outbound

This paper cites Exploration by random network distillation.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Exploration by random network distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.605429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.605429Z digest=sha256:c7893d00c83f3ee97a8885b3865c3e24b33453fd6344fc4953b055fec58cac18

Observation 4aeef1d3-1ded-45f7-8343-e31d037a09e1 · outbound

This paper cites Tighter bounds on the expressivity of transformer encoders.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Tighter bounds on the expressivity of transformer encoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.450779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.608516Z digest=sha256:871dfe7a99f8bb90c71cce6252679e30db72f66b0dc403901edd1b7bd2aed9c8

Observation f23df22b-6ad2-4664-b46d-00697d52e482 · outbound

This paper cites Metastability for general dynamics with rare transitions: escape time and critical configurations.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability for general dynamics with rare transitions: escape time and critical configurations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.441151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.611157Z digest=sha256:d240d99c36dd5cbb45cadeb40971cf791cb2c5d7ff4d9a844a85a1d834ed997a

Observation ec138d7e-1fef-4345-8cf3-21f1e2e22715 · outbound

This paper cites The Llama 3 Herd of Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.613806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.613806Z digest=sha256:45c817139e87fd1782398c1ef789e23e2ae4424c18dbb2522766a2a341f70e6d

Observation edff7053-f4c8-4965-865a-2b817574dfaf · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Edelman, eran malach, and Surbhi Goel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.431000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.616528Z digest=sha256:84b1b3d51ec0e390fbcb4fcdb077496387865539a2395e4e8ea29f5c2d390acd

Observation 3e56dd1e-f808-4a42-ba87-d0eac4056faf · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.421609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.619204Z digest=sha256:4a9f7cd80df081ed28afc992f144022752635264dc84df198350e5c5f042d540

Observation e2e99adf-8c49-403b-a03e-80512daef6c8 · outbound

This paper cites A general characterization of the statistical query complexity.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A general characterization of the statistical query complexity

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.411935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.621676Z digest=sha256:d3efd0a978c16db2998ea5f807c5068d842831cd769940de5634e21b8e2ea41b

Observation 89d70ac3-eb8a-4a3a-b481-e2b7952f8457 · outbound

This paper cites Towards revealing the mystery behind chain of thought: a theoretical perspective.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Towards revealing the mystery behind chain of thought: a theoretical perspective

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.401771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.624383Z digest=sha256:fe5cc77e1237dd67c5829ca1fc0508204d911ee4823a1989baf5c937f5472fbc

Observation 469e47ad-84b8-460e-a008-a747f6fe8abb · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.626962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.626962Z digest=sha256:901cf3913a5cdd7269753a94b3c79d37c040f7c0a0f112f0a2a2f8626da908cd

Observation 071849db-828d-4540-8e35-26bf1c123f34 · outbound

This paper cites Fernandez, F.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Fernandez, F

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.391558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.629858Z digest=sha256:1cf2ce72ce0286e77d5196117003efba8304d0a95cdfdc92ac627ededb7f0db4

Observation 7e8eda49-89cb-4d9e-8040-264ffc98f8e1 · outbound

This paper cites Asymptotically exponential hitting times and metastability: A pathwise approach without reversibility.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Asymptotically exponential hitting times and metastability: A pathwise approach without reversibility

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.382935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.632315Z digest=sha256:37359749a203843d0944c038046e9fe92255bd42eb6df6889e799fc09dc29965

Observation d0e991fa-b97f-4c3b-bcd4-5eeb67c1279d · outbound

This paper cites An SVD approach to identifying metastable states of Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation An SVD approach to identifying metastable states of Markov chains

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.374387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.635584Z digest=sha256:a87a73dd41db9678ddb0dba45a995551272ea5b9ee7593fbe38eeef3b88c3cca

Observation 48240f96-30af-43c8-8308-ccd706082de6 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Stream of Search (SoS): Learning to Search in Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.638911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.638911Z digest=sha256:5c86fafe0755d8ed660a19c5ddd1a8713c31b857e5e41298ed32fdbf3f38b747

Observation b92c176c-6045-4369-901f-0c9a3b86708b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.642300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.642300Z digest=sha256:10c2e65ad6c892d5fd25da298a93f86d4fe9d06e81a619aa3e3eca00cc40cf94

Observation dd70058d-78f5-45b0-a30b-b72a19a6cdbe · outbound

This paper cites Training Compute-Optimal Large Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Training Compute-Optimal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.645993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.645993Z digest=sha256:68892d116dcb92d586d32ade4d720ad9fb04fcd0ebc58ba81d9b94d26d910083

Observation e5b84f39-108a-4f7f-99bd-ec0a02abb40a · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.649626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.649626Z digest=sha256:1a6a46da2f4f45b09e8e1d16c4a257d9ca252c9cf0c9eabf318fb283396f646e

Observation 5a07c5ea-c178-4f41-9fb7-c849e8acde89 · outbound

This paper cites Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.653083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.653083Z digest=sha256:e0b5f2aef0bcb1662a3024e156320943bc27a52c68dc3dfb1fe4c25d5406970b

Observation b12e8f29-6e43-4430-ad5b-4bfd4ec8ae77 · outbound

This paper cites From self-attention to M arkov models: unveiling the dynamics of generative transformers.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation From self-attention to M arkov models: unveiling the dynamics of generative transformers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.365773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.656125Z digest=sha256:30302cd57ba17b00bfa03e45b79930e2b8ac04efcf04bad5ae5866499e2a879a

Observation 7afa93a0-6c1a-49e0-b280-44042fba37ea · outbound

This paper cites A robust spectral method for finding lumpings and meta-stable states of non-reversible M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A robust spectral method for finding lumpings and meta-stable states of non-reversible M arkov chains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.356263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.659523Z digest=sha256:7587d576e7e93f9c0fe6a0f7f93556c1ff32afe7bacbe897b7af5c05dbbe1250

Observation 37bf7a01-0b54-4bc9-b596-a91a54369a6f · outbound

This paper cites OpenAI o1 System Card.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation OpenAI o1 System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.662426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.662426Z digest=sha256:b7522fe2314d1036dbbde84cc18a28e6bcd9eb93ab1c501cd9b2aec2b413cce5

Observation e7498104-799f-4444-986d-865692d73cf5 · outbound

This paper cites Risk and parameter convergence of logistic regression.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Risk and parameter convergence of logistic regression

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.665817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.665817Z digest=sha256:906a3e6e8d8460b71e80f25cae58844ece8d178c846a3d74cc38a55dd44a050a

Observation 7fe2369e-b6da-4d2c-995e-96ac8a535b11 · outbound

This paper cites Scaling Scaling Laws with Board Games.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling Scaling Laws with Board Games

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.669105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.669105Z digest=sha256:a574adf72627f6310116e02abec7e0f8774a9efd457c6521008a5159dde31307

Observation 2417e27c-3988-47b8-b0f1-fe2d0c01df37 · outbound

This paper cites Thinking, fast and slow.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Thinking, fast and slow

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.672379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.672379Z digest=sha256:230784cc95323ff75b76cd4bbcebacaa742271bf563db1d461063bd5830c9d93

Observation b6918358-61f6-47a2-a91f-5ff3a45743a7 · outbound

This paper cites Scaling Laws for Neural Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.675692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.675692Z digest=sha256:eb46a0a0b0146f4dbad36ebedf47270a737cd0cb3c5dec85b3e91694160380d0

Observation c345ce1e-e0f3-40ec-9b08-2b61dd640948 · outbound

This paper cites Efficient noise-tolerant learning from statistical queries.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Efficient noise-tolerant learning from statistical queries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.341073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.678892Z digest=sha256:cf7af64a3c2cb33707e9bb2fab5a1334e47570fd3f8fb62c46613a4152db0860

Observation 5faf387d-a41d-4915-a41a-42faf968e9ea · outbound

This paper cites Transformers Provably Solve Parity Efficiently with Chain of Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Transformers Provably Solve Parity Efficiently with Chain of Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.682061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.682061Z digest=sha256:dfadaa3b055578cfbfb278f5274382f641cbf9ee50252775ced274de1379989f

Observation 30225a18-51ae-4175-88cd-af4693d0674d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.685445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.685445Z digest=sha256:5d266bb47d5953c7e62115b749b0da77d8273d5635d906496bb59818ea431238

Observation 8b264acb-2949-47cf-9403-c39bc63f3a80 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Training Language Models to Self-Correct via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.688957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.688957Z digest=sha256:f9170008d486ed3f688fbe735b92086d8026f1ca281af0bd7e8b3dbbfddcae41

Observation 08acdfda-e33b-48b1-847e-7e9b986c704b · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.331402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.692341Z digest=sha256:80df2435d765b6eee95e1afcf9fc03e0501e4889501c0648e62724765bd228c7

Observation 03e0c432-004e-4585-8ef0-1c9ef6347810 · outbound

This paper cites Metastable Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastable Markov chains

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:41.020448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.695286Z digest=sha256:e80c124725f5ca1bc53acc190c9008af0a69e0f28e890f19e2c80c4d504e0979

Observation 12268893-11f3-41f1-8e37-ac1dcf0009e5 · outbound

This paper cites Metastability of finite state M arkov chains: A recursive procedure to identify slow variables for model reduction.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability of finite state M arkov chains: A recursive procedure to identify slow variables for model reduction

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.322027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.698504Z digest=sha256:4a8a96bddf75428b6a9cf7ee0571aa9d96c09d0d6866d2956a6a46d32a8761c9

Observation 77b3410a-cd18-4ff9-8885-affb135f9885 · outbound

This paper cites Markov Chains and Mixing Times.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Markov Chains and Mixing Times

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.312481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.701182Z digest=sha256:a838bf5ebdd911783fe2789c2f295f632b9f94e6676e486a4784c0fab92c3221

Observation 2f7d99d4-2f58-4def-a562-a386a6d5ff2c · outbound

This paper cites How do nonlinear transformers acquire generalization-guaranteed CoT ability? In High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning, 2024 a.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How do nonlinear transformers acquire generalization-guaranteed CoT ability? In High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning, 2024 a

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.303024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.703758Z digest=sha256:442b5cd127ca72bcc1b639e15ce222a28082e4f89f86cf13bd8278a425ca38e1

Observation 5d094478-9b26-4cf9-99d2-26cd5fd48885 · outbound

This paper cites Dissecting chain-of-thought: compositionality through in-context filtering and learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Dissecting chain-of-thought: compositionality through in-context filtering and learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.293444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.706269Z digest=sha256:f9815467f78e4025111126fd0e722ceb95ea857d0b09661496fef47adee9cf02

Observation b353fa84-b2d0-4ea2-ab11-22513e1d936e · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.708818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.708818Z digest=sha256:cb3182ba2a0cf143bddea3f7df20b39e680fe4671ea61cba885a3f0b59155718

Observation 24774810-46a9-4268-9075-14a390b6444b · outbound

This paper cites Let's Verify Step by Step.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Let's Verify Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.711533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.711533Z digest=sha256:2267201c801e227bd13bb6d232df22f211e85b8735a57f719e81d7647ef673d7

Observation 9961a831-396c-4a27-a69d-1b8d65c0a58e · outbound

This paper cites Markov chain decomposition for convergence rate analysis.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Markov chain decomposition for convergence rate analysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.283333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.714354Z digest=sha256:5b4dfccb0daf5c0e6900a8fedf60be7776249f07d3c2fe9c36a5c335f63f41a2

Observation 5c5eceea-8ea7-477f-87b3-9750e7750afd · outbound

This paper cites Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.717508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.717508Z digest=sha256:02188f4c71f01973785a7b1861d0551ad9ba02f7e8d88107af61f48da302f2b1

Observation ac800233-a78a-4198-a1c7-f973b4176e20 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The Expressive Power of Transformers with Chain of Thought

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.720858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.720858Z digest=sha256:0144fa54cf42f3a8daebf61335e540c4c162f43b14ec4e210c9aa677df32eb50

Observation 65e29515-9988-4385-95a3-0bdd90c367e6 · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.274007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.724628Z digest=sha256:c34cfe87c7b7a9a0f1a86c599d517f94bc4ae275925b0f5fe79abf5d906108be

Observation 995399be-ef35-42b2-8eba-547f5edfb5ee · outbound

This paper cites The dondition of a finite M arkov chain and perturbation bounds for the limiting probabilities.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The dondition of a finite M arkov chain and perturbation bounds for the limiting probabilities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.264868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.727588Z digest=sha256:e3f7858cf112f8e0425565696243db4afa414595f253bd7654fa49147e0be2d1

Observation b4b8a0cc-5aff-4b01-91e0-42c87d69b3f1 · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How Transformers Learn Causal Structure with Gradient Descent

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.730834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.730834Z digest=sha256:f80e7898acc9fa30232ff4a224d688573a3015c788d4fe05c28d25a24a8e9371

Observation 308c9202-df0e-4a94-b4cc-e293a707d617 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.734065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.734065Z digest=sha256:c7f768bac55489425f0bb6f0570f198fbac266ca4ebd43443a5cbd3a8c371762

Observation b7ab6864-ae78-42e4-bb54-d28681e259c2 · outbound

This paper cites Spinning up: proximal policy optimization ( PPO ), 2018.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Spinning up: proximal policy optimization ( PPO ), 2018

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.255318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.737533Z digest=sha256:5cb59f0d35cc038f6a8e8076ab5be3281b910b6544e5815f82ad8b9104f7c30f

Observation f4159ad9-0374-4bf2-a78f-3e6e8c63a2d1 · outbound

This paper cites Concentration inequalities for Markov chains by Marton couplings and spectral methods.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Concentration inequalities for Markov chains by Marton couplings and spectral methods

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.245878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.740784Z digest=sha256:581f140e08e1f9bfe10971c12efe83483ec327b87abc8918680d8aa29a6341a7

Observation 744c92bc-a181-4673-ac27-48d0f55c058c · outbound

This paper cites Improving language understanding by generative pre-training.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Improving language understanding by generative pre-training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.743850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.743850Z digest=sha256:ae755162d149db220bc0118288982122ae7b2fd4dff47661b11c396a5fc7c677

Observation 2669e4e7-1f0e-4ab0-8527-94308893de45 · outbound

This paper cites Understanding Transformer Reasoning Capabilities via Graph Algorithms.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Understanding Transformer Reasoning Capabilities via Graph Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.747108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.747108Z digest=sha256:e35b5ccaf32e9ddd91859f081504873c3b82970e488bb984ceb9efd8dd755fa3

Observation 9a4f38dc-8c81-4cbb-a9c8-51502a1928b1 · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Transformers, parallel computation, and logarithmic depth

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.231676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.750617Z digest=sha256:38d4bd43491cb2218ef93c094734da6b4eb49acca5f5acfe4b0e3b01187f4845

Observation 968d5237-196b-42c2-9c15-56e5adc7d758 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Proximal Policy Optimization Algorithms

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.753852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.753852Z digest=sha256:1fe08918a0c876a49f1579ec082fc560e94fa4da8ff6f90fca40c8f687febb13

Observation 61f1fcab-3a01-473e-a46b-d9ff406aabc6 · outbound

This paper cites Failures of gradient-based deep learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Failures of gradient-based deep learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.223371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.757299Z digest=sha256:af161870d18de69b2dd4932787774931150d557939db41b00638bfcf7f5bdb10

Observation 91772f9d-6a93-4736-8c97-60796dc9d278 · outbound

This paper cites Distribution-specific hardness of learning neural networks.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distribution-specific hardness of learning neural networks

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.214982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.760427Z digest=sha256:6d12400ef691125faad0841ed2c0c1dab4bfa2495d3d7c756b4ccc39af7958a4

Observation c7785893-e39f-48a0-bc0d-a4e16d8f8e84 · outbound

This paper cites Distilling Reasoning Capabilities into Smaller Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distilling Reasoning Capabilities into Smaller Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.763610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.763610Z digest=sha256:e7f62c2b3eaf3198aaeb0be7bb9fbd1a8dcffb7ea83bffc6afb484355f61a270

Observation 296fcd87-0aae-4d82-af24-2e38d886710c · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and G o through self-play.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A general reinforcement learning algorithm that masters chess, shogi, and G o through self-play

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.205822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.766915Z digest=sha256:6cd158590beb21c71fc2647611c94421917f89bfe6ee86c76e2b78f8217ce7d3

Observation 5bb19291-011c-4e56-bafb-bb65cecde4fe · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.770120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.770120Z digest=sha256:28ef780b569f3519989a6fc256f8999bb50c6ff4f086be9b60efb0b8be62a073

Observation 000de0fa-d41b-4a9d-8255-83c582ff224b · outbound

This paper cites On an SVD-based algorithm for identifying meta-stable states of Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation On an SVD-based algorithm for identifying meta-stable states of Markov chains

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.196240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.773565Z digest=sha256:adbf7a1fd17e319f28e6035c43e2388de78e5ea07d0cfdaef675bebd506dcb02

Observation 8166e16f-0343-40fe-806b-e150555df2f0 · outbound

This paper cites Solving olympiad geometry without human demonstrations.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Solving olympiad geometry without human demonstrations

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.777217Z digest=sha256:1e0d8996984043f1079f992f67f2682eb6e127739faf64b41740b2721dbde47f

Observation ad2d2566-6750-4f48-9ff7-74600fb4dbbc · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Solving math word problems with process- and outcome-based feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.780527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.780527Z digest=sha256:be7c3d9acc572c551f6cd17f771af9e1f408d725969f5d9979d0d15755652778

Observation f90ef72d-ac50-4ee3-8860-8b9b19b37220 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.784127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.784127Z digest=sha256:87f3b2bf9bad89c16328d84a79e4cb402dca1bea5a64f88c17f509f0507b4412

Observation 7dda2a9b-e9a1-410b-b2db-4868a9765058 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.787506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.787506Z digest=sha256:3a6cd01fee551f70dfb5ed2aca204bada7cc1f3a1759ebfa15a35a66f46cddd5

Observation 0c50f4c7-5063-4dec-9504-5d105b643444 · outbound

This paper cites An algorithm for computing stochastically stable distributions with applications to mltiagent learning in repeated games.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation An algorithm for computing stochastically stable distributions with applications to mltiagent learning in repeated games

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.175635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.790869Z digest=sha256:b19d6ff0cc49282218e930c20b887a9f961c2f9728848c6521f57dd50339077b

Observation 3a93b245-11d4-4632-909d-d86fde763e0a · outbound

This paper cites Estimating the Mixing Time of Ergodic Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Estimating the Mixing Time of Ergodic Markov Chains

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:40.897167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.793651Z digest=sha256:2031d361ab4c0c491b39358b43c518041e720e97a41f88cf32ce3c81b97daa2c

Observation f916fd35-a510-4a5d-a42b-d1c159d687b9 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.796752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.796752Z digest=sha256:cb1b3e3c389389ce99bc6c55e1ad567c3caa86b1c24d940ae13e81b09068cdea

Observation 1d7d5871-f5a4-48c6-919a-71dabc5f150f · outbound

This paper cites Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.799741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.799741Z digest=sha256:0b430b6a27d33c65821281e28e027e85d0f327ef38c89ed7cb34283f068c3692

Observation 77bd2389-0d2d-41bf-8f95-2dc1e2df1484 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.802953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.802953Z digest=sha256:5e8721f1a6e85010026c77ae5e18379d825c55625f0b31e4e3f0798c11f5fafe

Observation 8e536755-9a40-4a33-93ff-19d5b89a3da3 · outbound

This paper cites What Can Neural Networks Reason About?.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation What Can Neural Networks Reason About?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.806200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.806200Z digest=sha256:049a42a5703a228b551dd02540c91fb26c2eae4dfa5a97db69cb83dce98321ea

Observation 65f680de-6147-4b82-9b70-1f33cafdff9b · outbound

This paper cites Tree of thoughts: deliberate problem solving with large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Tree of thoughts: deliberate problem solving with large language models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.809759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.809759Z digest=sha256:7b83201c8ff5d671dde408b49bd1d4f779e2bb7925b6b7837187a25b60b19a45

Observation 4eb8bfad-a55f-4ce3-9f4f-03d7de9d50fc · outbound

This paper cites Large Language Models as Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Large Language Models as Markov Chains

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.812795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.812795Z digest=sha256:3d2911832865e23fdb9b2f437c8405d52071c28d1735d3a92733e90cd617288e

Observation c58d727f-2cd0-43b7-af55-33038c0db0f3 · outbound

This paper cites Star: bootstrapping reasoning with reasoning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Star: bootstrapping reasoning with reasoning

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.160931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.815997Z digest=sha256:c1b1493e07a8c86f74bd7253119b0954429503ce74b2315bafcbfeaed36a54db

Pith citing papers

Observation 8c284df7-dd23-469d-8630-b8d80b83a028 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

Reference 300

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:23.937668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:81dd5c017afe10391a025f8f81a26c90bfb08296628767ec586ebafa1a088b46

Observation 2cd0d59c-5a84-4693-8361-82fbf94ab6ec · inbound

The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently cites this paper.

The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:24.330367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T05:27:11.761971Z digest=sha256:f3d86120a61a290ef76440ea97fca17be7a72308ac1a031bfb58b3fbde98d02e