Pith. sign in

Paper Citation Record · LEDGER

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

As of 10 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 2 inbound Pith citation observations for arXiv:2502.01694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01694 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:40.815997Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:36:23.845366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:36:23.932030Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50c00ae1-3b43-4ffa-b4e7-ba111481a8b4 · outbound

This paper cites How far can transformers reason? T he locality barrier and inductive scratchpad.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How far can transformers reason? T he locality barrier and inductive scratchpad

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.536637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.569761Z digest=sha256:60702bd535cee837703d0b8f87390347a49d8b6dc88966689fda846b4ddcac9e

Observation e438d207-4bb6-4b9f-896e-5d1448966b19 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.574169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.574169Z digest=sha256:bdc64221ea498487f0a4b0b70f2f6558cbf9aa8290fa42542b26129a3f0d29f2

Observation 24a5fb29-86a9-4465-b1a8-8593739dab93 · outbound

This paper cites Beltr\' a n and C.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Beltr\' a n and C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.528255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.578030Z digest=sha256:11c9089fa6320a6fdf054a560a0ae0fc267212d647f61f1484316bac41affd38

Observation ac51a1ab-8cde-4ef2-b566-066b90401c98 · outbound

This paper cites Graph of thoughts: solving elaborate problems with large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Graph of thoughts: solving elaborate problems with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.520267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.581505Z digest=sha256:796ed8c387850245014576a494f926d0a697c077b04aea00fb34d290a92fa1a8

Observation 0cb9b975-6509-42de-b57d-622616770bd3 · outbound

This paper cites Multi-scale metastable dynamics and the asymptotic stationary distribution of perturbed M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Multi-scale metastable dynamics and the asymptotic stationary distribution of perturbed M arkov chains

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.511374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.585137Z digest=sha256:eb21c8eebf7a86964cadd6beed897e7680132016e746b848a79fab45b15c3867

Observation 2b994676-eff7-4a6f-8c49-7003c4e6a366 · outbound

This paper cites Understanding in-context learning in transformers and LLM s by learning to learn discrete functions.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Understanding in-context learning in transformers and LLM s by learning to learn discrete functions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.588486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.588486Z digest=sha256:87f3aa26d39688dcaf2d435a7d1b1d7f2bc35ec94e775fdc088ed417901de245

Observation 6c81d042-285a-44ee-b136-05f8e29d1f98 · outbound

This paper cites Metastable states, quasi-stationary distributions and soft measures.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastable states, quasi-stationary distributions and soft measures

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.495805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.592115Z digest=sha256:cd20e23b617949dcade933e9bef267ebfd1a260a6376d5a227a8036a3686109e

Observation 89100101-25d3-4c9c-86e4-ec4e3829ef17 · outbound

This paper cites Metastability and low lying spectra in reversible M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability and low lying spectra in reversible M arkov chains

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.486093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.595430Z digest=sha256:775eb7b13e72794e3d44f1d9515aea8105231b591ac41822ef59e63c08dc2491

Observation 9dd31d78-eb4b-4ea7-b163-a0b46229129f · outbound

This paper cites On using extended statistical queries to avoid membership queries.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation On using extended statistical queries to avoid membership queries

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.476089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.598732Z digest=sha256:ecf471bfd9be162d8468bc681d9f13bf0dbb660d7e01f521b859c3733c2eea06

Observation 54e4cb9c-3218-473e-a04b-8ab501f75d64 · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.466392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.602193Z digest=sha256:2a2fbe1dfe470f300023524043091f8ac58342aaa4d3ba901467af07e42b1273

Observation 19df1bd3-babe-439a-aa57-3fb4c06146de · outbound

This paper cites Exploration by random network distillation.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Exploration by random network distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.605429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.605429Z digest=sha256:c7893d00c83f3ee97a8885b3865c3e24b33453fd6344fc4953b055fec58cac18

Observation 4aeef1d3-1ded-45f7-8343-e31d037a09e1 · outbound

This paper cites Tighter bounds on the expressivity of transformer encoders.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Tighter bounds on the expressivity of transformer encoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.450779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.608516Z digest=sha256:ad4b8dd2ec9c5065d28e7c205b3f02bada5e94032a7d6b58ee093598a5181062

Observation f23df22b-6ad2-4664-b46d-00697d52e482 · outbound

This paper cites Metastability for general dynamics with rare transitions: escape time and critical configurations.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability for general dynamics with rare transitions: escape time and critical configurations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.441151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.611157Z digest=sha256:5210187a690e153dfd32a4b1c9b2a632f49b3011b92fa33af308d2aa2cebb00d

Observation ec138d7e-1fef-4345-8cf3-21f1e2e22715 · outbound

This paper cites The Llama 3 Herd of Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.613806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.613806Z digest=sha256:45c817139e87fd1782398c1ef789e23e2ae4424c18dbb2522766a2a341f70e6d

Observation edff7053-f4c8-4965-865a-2b817574dfaf · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Edelman, eran malach, and Surbhi Goel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.431000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.616528Z digest=sha256:7e213050340cd3ef644757356382ba3ceaade73175177abb0a093b5557ef7675

Observation 3e56dd1e-f808-4a42-ba87-d0eac4056faf · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.421609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.619204Z digest=sha256:7085f183f7be8c093c863c4e9086d08bb209babe5eda7e5e32d3eefdfa15264c

Observation e2e99adf-8c49-403b-a03e-80512daef6c8 · outbound

This paper cites A general characterization of the statistical query complexity.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A general characterization of the statistical query complexity

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.411935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.621676Z digest=sha256:9be7c7353b645e9f6a2bce021d47c5f53f8d9a74b428e98226d12ec87bfedf31

Observation 89d70ac3-eb8a-4a3a-b481-e2b7952f8457 · outbound

This paper cites Towards revealing the mystery behind chain of thought: a theoretical perspective.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Towards revealing the mystery behind chain of thought: a theoretical perspective

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.401771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.624383Z digest=sha256:afd06f09d931d0bd9c6c6af523d27e12e69a237c87a1e8b5e9fda3f98a6cdc0a

Observation 469e47ad-84b8-460e-a008-a747f6fe8abb · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.626962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.626962Z digest=sha256:901cf3913a5cdd7269753a94b3c79d37c040f7c0a0f112f0a2a2f8626da908cd

Observation 071849db-828d-4540-8e35-26bf1c123f34 · outbound

This paper cites Fernandez, F.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Fernandez, F

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.391558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.629858Z digest=sha256:d80ed2c201effe8ac882aa3f610d0d44685d5617c9734a30dafde1e7a1f72ace

Observation 7e8eda49-89cb-4d9e-8040-264ffc98f8e1 · outbound

This paper cites Asymptotically exponential hitting times and metastability: A pathwise approach without reversibility.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Asymptotically exponential hitting times and metastability: A pathwise approach without reversibility

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.382935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.632315Z digest=sha256:a8918cdd76044cbba6db7237d4d1f75e60ca12fffdb32455436ecece45e26698

Observation d0e991fa-b97f-4c3b-bcd4-5eeb67c1279d · outbound

This paper cites An SVD approach to identifying metastable states of Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation An SVD approach to identifying metastable states of Markov chains

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.374387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.635584Z digest=sha256:0da878b74b0a97c9cc30421fcf53007ebdde6f06d5f9532e3d0546be5ebacf31

Observation 48240f96-30af-43c8-8308-ccd706082de6 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Stream of Search (SoS): Learning to Search in Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.638911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.638911Z digest=sha256:5c86fafe0755d8ed660a19c5ddd1a8713c31b857e5e41298ed32fdbf3f38b747

Observation b92c176c-6045-4369-901f-0c9a3b86708b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.642300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.642300Z digest=sha256:10c2e65ad6c892d5fd25da298a93f86d4fe9d06e81a619aa3e3eca00cc40cf94

Observation dd70058d-78f5-45b0-a30b-b72a19a6cdbe · outbound

This paper cites Training Compute-Optimal Large Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Training Compute-Optimal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.645993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.645993Z digest=sha256:68892d116dcb92d586d32ade4d720ad9fb04fcd0ebc58ba81d9b94d26d910083

Observation e5b84f39-108a-4f7f-99bd-ec0a02abb40a · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.649626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.649626Z digest=sha256:1a6a46da2f4f45b09e8e1d16c4a257d9ca252c9cf0c9eabf318fb283396f646e

Observation 5a07c5ea-c178-4f41-9fb7-c849e8acde89 · outbound

This paper cites Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.653083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.653083Z digest=sha256:e0b5f2aef0bcb1662a3024e156320943bc27a52c68dc3dfb1fe4c25d5406970b

Observation b12e8f29-6e43-4430-ad5b-4bfd4ec8ae77 · outbound

This paper cites From self-attention to M arkov models: unveiling the dynamics of generative transformers.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation From self-attention to M arkov models: unveiling the dynamics of generative transformers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.365773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.656125Z digest=sha256:4e68d8ea6fb1f127cdb93d831164db164b578dda049eac1deb55460584f6ad11

Observation 7afa93a0-6c1a-49e0-b280-44042fba37ea · outbound

This paper cites A robust spectral method for finding lumpings and meta-stable states of non-reversible M arkov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A robust spectral method for finding lumpings and meta-stable states of non-reversible M arkov chains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.356263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.659523Z digest=sha256:4715319d116723c903f651d81a3e4052adb92dd1f948739d91e2deffb800f894

Observation 37bf7a01-0b54-4bc9-b596-a91a54369a6f · outbound

This paper cites OpenAI o1 System Card.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation OpenAI o1 System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.662426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.662426Z digest=sha256:b7522fe2314d1036dbbde84cc18a28e6bcd9eb93ab1c501cd9b2aec2b413cce5

Observation e7498104-799f-4444-986d-865692d73cf5 · outbound

This paper cites Risk and parameter convergence of logistic regression.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Risk and parameter convergence of logistic regression

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.665817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.665817Z digest=sha256:906a3e6e8d8460b71e80f25cae58844ece8d178c846a3d74cc38a55dd44a050a

Observation 7fe2369e-b6da-4d2c-995e-96ac8a535b11 · outbound

This paper cites Scaling Scaling Laws with Board Games.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling Scaling Laws with Board Games

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.669105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.669105Z digest=sha256:a574adf72627f6310116e02abec7e0f8774a9efd457c6521008a5159dde31307

Observation 2417e27c-3988-47b8-b0f1-fe2d0c01df37 · outbound

This paper cites Thinking, fast and slow.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Thinking, fast and slow

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.672379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.672379Z digest=sha256:230784cc95323ff75b76cd4bbcebacaa742271bf563db1d461063bd5830c9d93

Observation b6918358-61f6-47a2-a91f-5ff3a45743a7 · outbound

This paper cites Scaling Laws for Neural Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.675692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.675692Z digest=sha256:eb46a0a0b0146f4dbad36ebedf47270a737cd0cb3c5dec85b3e91694160380d0

Observation c345ce1e-e0f3-40ec-9b08-2b61dd640948 · outbound

This paper cites Efficient noise-tolerant learning from statistical queries.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Efficient noise-tolerant learning from statistical queries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.341073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.678892Z digest=sha256:bbec5f781715f110be80a89ed967bada54416f15eabd692d160db0d4d4a39dca

Observation 5faf387d-a41d-4915-a41a-42faf968e9ea · outbound

This paper cites Transformers Provably Solve Parity Efficiently with Chain of Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Transformers Provably Solve Parity Efficiently with Chain of Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.682061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.682061Z digest=sha256:dfadaa3b055578cfbfb278f5274382f641cbf9ee50252775ced274de1379989f

Observation 30225a18-51ae-4175-88cd-af4693d0674d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.685445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.685445Z digest=sha256:5d266bb47d5953c7e62115b749b0da77d8273d5635d906496bb59818ea431238

Observation 8b264acb-2949-47cf-9403-c39bc63f3a80 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Training Language Models to Self-Correct via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.688957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.688957Z digest=sha256:f9170008d486ed3f688fbe735b92086d8026f1ca281af0bd7e8b3dbbfddcae41

Observation 08acdfda-e33b-48b1-847e-7e9b986c704b · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.331402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.692341Z digest=sha256:1b895ca1315bace94dd2aa45884bd7bfa09570cf650153a141fa5d1673684ab8

Observation 03e0c432-004e-4585-8ef0-1c9ef6347810 · outbound

This paper cites Metastable Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastable Markov chains

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:41.020448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.695286Z digest=sha256:3c1ede77f53c78f6b3f3247dcfd40bcad9d92603af21f50388c4e215777aff86

Observation 12268893-11f3-41f1-8e37-ac1dcf0009e5 · outbound

This paper cites Metastability of finite state M arkov chains: A recursive procedure to identify slow variables for model reduction.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Metastability of finite state M arkov chains: A recursive procedure to identify slow variables for model reduction

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.322027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.698504Z digest=sha256:05d8e2c494f481fabc296308a29fe794bc4f52fbef6780eb23a61fd72aa16e0e

Observation 77b3410a-cd18-4ff9-8885-affb135f9885 · outbound

This paper cites Markov Chains and Mixing Times.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Markov Chains and Mixing Times

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.312481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.701182Z digest=sha256:b7a575c6ef9ecfaa251dc7950a2b50bdc63b4841a5a64a2c52a2573523fe6bdf

Observation 2f7d99d4-2f58-4def-a562-a386a6d5ff2c · outbound

This paper cites How do nonlinear transformers acquire generalization-guaranteed CoT ability? In High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning, 2024 a.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How do nonlinear transformers acquire generalization-guaranteed CoT ability? In High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning, 2024 a

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.303024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.703758Z digest=sha256:bd2b3082c61fac1d42cd5e9febe4bd982b404975fcfd74b0918238c469795546

Observation 5d094478-9b26-4cf9-99d2-26cd5fd48885 · outbound

This paper cites Dissecting chain-of-thought: compositionality through in-context filtering and learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Dissecting chain-of-thought: compositionality through in-context filtering and learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.293444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.706269Z digest=sha256:eb94ca2b057a037a287d7a69ff2b0b6407faf3c1e4322a39e10e77f81195cfb0

Observation b353fa84-b2d0-4ea2-ab11-22513e1d936e · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.708818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.708818Z digest=sha256:cb3182ba2a0cf143bddea3f7df20b39e680fe4671ea61cba885a3f0b59155718

Observation 24774810-46a9-4268-9075-14a390b6444b · outbound

This paper cites Let's Verify Step by Step.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Let's Verify Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.711533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.711533Z digest=sha256:2267201c801e227bd13bb6d232df22f211e85b8735a57f719e81d7647ef673d7

Observation 9961a831-396c-4a27-a69d-1b8d65c0a58e · outbound

This paper cites Markov chain decomposition for convergence rate analysis.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Markov chain decomposition for convergence rate analysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.283333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.714354Z digest=sha256:85c25a262421789cadf2fd1f009b4733ab412a27e94b8ca2b313ff2c3f407eda

Observation 5c5eceea-8ea7-477f-87b3-9750e7750afd · outbound

This paper cites Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.717508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.717508Z digest=sha256:02188f4c71f01973785a7b1861d0551ad9ba02f7e8d88107af61f48da302f2b1

Observation ac800233-a78a-4198-a1c7-f973b4176e20 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The Expressive Power of Transformers with Chain of Thought

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.720858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.720858Z digest=sha256:0144fa54cf42f3a8daebf61335e540c4c162f43b14ec4e210c9aa677df32eb50

Observation 65e29515-9988-4385-95a3-0bdd90c367e6 · outbound

This paper cites an unresolved cited work.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:41.274007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.724628Z digest=sha256:8f4bbfe6b04948480184cb9abeee8e1d2d926f370ff56b2f7b7447ae59e955cb

Observation 995399be-ef35-42b2-8eba-547f5edfb5ee · outbound

This paper cites The dondition of a finite M arkov chain and perturbation bounds for the limiting probabilities.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation The dondition of a finite M arkov chain and perturbation bounds for the limiting probabilities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.264868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.727588Z digest=sha256:e9e38bf79f77023ca2cf7ecb91a733dc0b7f24db20f454723923e4cebef0359b

Observation b4b8a0cc-5aff-4b01-91e0-42c87d69b3f1 · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How Transformers Learn Causal Structure with Gradient Descent

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.730834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.730834Z digest=sha256:f80e7898acc9fa30232ff4a224d688573a3015c788d4fe05c28d25a24a8e9371

Observation 308c9202-df0e-4a94-b4cc-e293a707d617 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.734065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.734065Z digest=sha256:c7f768bac55489425f0bb6f0570f198fbac266ca4ebd43443a5cbd3a8c371762

Observation b7ab6864-ae78-42e4-bb54-d28681e259c2 · outbound

This paper cites Spinning up: proximal policy optimization ( PPO ), 2018.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Spinning up: proximal policy optimization ( PPO ), 2018

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.255318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.737533Z digest=sha256:44cc665af84d4fbdcb43860ee44257dcbca8566d136ef2d22643b1fd8bdfe1c3

Observation f4159ad9-0374-4bf2-a78f-3e6e8c63a2d1 · outbound

This paper cites Concentration inequalities for Markov chains by Marton couplings and spectral methods.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Concentration inequalities for Markov chains by Marton couplings and spectral methods

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.245878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.740784Z digest=sha256:dfd89d40c44e2a4fee98cf4af53573e96cb7b85c6813f38f6dbb26a1fe1a7dce

Observation 744c92bc-a181-4673-ac27-48d0f55c058c · outbound

This paper cites Improving language understanding by generative pre-training.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Improving language understanding by generative pre-training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.743850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.743850Z digest=sha256:ae755162d149db220bc0118288982122ae7b2fd4dff47661b11c396a5fc7c677

Observation 2669e4e7-1f0e-4ab0-8527-94308893de45 · outbound

This paper cites Understanding Transformer Reasoning Capabilities via Graph Algorithms.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Understanding Transformer Reasoning Capabilities via Graph Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.747108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.747108Z digest=sha256:e35b5ccaf32e9ddd91859f081504873c3b82970e488bb984ceb9efd8dd755fa3

Observation 9a4f38dc-8c81-4cbb-a9c8-51502a1928b1 · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Transformers, parallel computation, and logarithmic depth

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.231676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.750617Z digest=sha256:b87cdf1f88647418ffd8fd41a75489d5d2041c300850558ca9afeea9580479dc

Observation 968d5237-196b-42c2-9c15-56e5adc7d758 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Proximal Policy Optimization Algorithms

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.753852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.753852Z digest=sha256:1fe08918a0c876a49f1579ec082fc560e94fa4da8ff6f90fca40c8f687febb13

Observation 61f1fcab-3a01-473e-a46b-d9ff406aabc6 · outbound

This paper cites Failures of gradient-based deep learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Failures of gradient-based deep learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.223371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.757299Z digest=sha256:7d6bb98d603c7afc4e64ae86742961c461621568b717569f7e58c88a307871a1

Observation 91772f9d-6a93-4736-8c97-60796dc9d278 · outbound

This paper cites Distribution-specific hardness of learning neural networks.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distribution-specific hardness of learning neural networks

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.214982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.760427Z digest=sha256:116ae37af86faaec1f658f02dac2d292424dfedded7028fb9a2a271a88c8a53b

Observation c7785893-e39f-48a0-bc0d-a4e16d8f8e84 · outbound

This paper cites Distilling Reasoning Capabilities into Smaller Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Distilling Reasoning Capabilities into Smaller Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.763610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.763610Z digest=sha256:e7f62c2b3eaf3198aaeb0be7bb9fbd1a8dcffb7ea83bffc6afb484355f61a270

Observation 296fcd87-0aae-4d82-af24-2e38d886710c · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and G o through self-play.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation A general reinforcement learning algorithm that masters chess, shogi, and G o through self-play

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.205822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.766915Z digest=sha256:854ce5b55b7db9552e61d8d7213df283243959c30edf11b563b2522b01393495

Observation 5bb19291-011c-4e56-bafb-bb65cecde4fe · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.770120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.770120Z digest=sha256:28ef780b569f3519989a6fc256f8999bb50c6ff4f086be9b60efb0b8be62a073

Observation 000de0fa-d41b-4a9d-8255-83c582ff224b · outbound

This paper cites On an SVD-based algorithm for identifying meta-stable states of Markov chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation On an SVD-based algorithm for identifying meta-stable states of Markov chains

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.196240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.773565Z digest=sha256:ca94993c2f38039753c8a13329d87c8d8d16e23bb719d5e6c41de51825bfbafb

Observation 8166e16f-0343-40fe-806b-e150555df2f0 · outbound

This paper cites Solving olympiad geometry without human demonstrations.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Solving olympiad geometry without human demonstrations

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.777217Z digest=sha256:1e0d8996984043f1079f992f67f2682eb6e127739faf64b41740b2721dbde47f

Observation ad2d2566-6750-4f48-9ff7-74600fb4dbbc · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Solving math word problems with process- and outcome-based feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.780527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.780527Z digest=sha256:be7c3d9acc572c551f6cd17f771af9e1f408d725969f5d9979d0d15755652778

Observation f90ef72d-ac50-4ee3-8860-8b9b19b37220 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.784127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.784127Z digest=sha256:87f3b2bf9bad89c16328d84a79e4cb402dca1bea5a64f88c17f509f0507b4412

Observation 7dda2a9b-e9a1-410b-b2db-4868a9765058 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.787506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.787506Z digest=sha256:3a6cd01fee551f70dfb5ed2aca204bada7cc1f3a1759ebfa15a35a66f46cddd5

Observation 0c50f4c7-5063-4dec-9504-5d105b643444 · outbound

This paper cites An algorithm for computing stochastically stable distributions with applications to mltiagent learning in repeated games.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation An algorithm for computing stochastically stable distributions with applications to mltiagent learning in repeated games

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.175635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.790869Z digest=sha256:2b34da4b14871a54e844d24632d0bbf22995c5d3a8eb60eaffe930c3544c2ef2

Observation 3a93b245-11d4-4632-909d-d86fde763e0a · outbound

This paper cites Estimating the Mixing Time of Ergodic Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Estimating the Mixing Time of Ergodic Markov Chains

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:40.897167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.793651Z digest=sha256:251b1cfe39bb711fff41246a15160a77f4e572cfb17e91c20e8ff43f82819ad2

Observation f916fd35-a510-4a5d-a42b-d1c159d687b9 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.796752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.796752Z digest=sha256:cb1b3e3c389389ce99bc6c55e1ad567c3caa86b1c24d940ae13e81b09068cdea

Observation 1d7d5871-f5a4-48c6-919a-71dabc5f150f · outbound

This paper cites Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.799741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.799741Z digest=sha256:0b430b6a27d33c65821281e28e027e85d0f327ef38c89ed7cb34283f068c3692

Observation 77bd2389-0d2d-41bf-8f95-2dc1e2df1484 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.802953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.802953Z digest=sha256:5e8721f1a6e85010026c77ae5e18379d825c55625f0b31e4e3f0798c11f5fafe

Observation 8e536755-9a40-4a33-93ff-19d5b89a3da3 · outbound

This paper cites What Can Neural Networks Reason About?.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation What Can Neural Networks Reason About?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.806200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.806200Z digest=sha256:049a42a5703a228b551dd02540c91fb26c2eae4dfa5a97db69cb83dce98321ea

Observation 65f680de-6147-4b82-9b70-1f33cafdff9b · outbound

This paper cites Tree of thoughts: deliberate problem solving with large language models.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Tree of thoughts: deliberate problem solving with large language models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.809759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.809759Z digest=sha256:7b83201c8ff5d671dde408b49bd1d4f779e2bb7925b6b7837187a25b60b19a45

Observation 4eb8bfad-a55f-4ce3-9f4f-03d7de9d50fc · outbound

This paper cites Large Language Models as Markov Chains.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Large Language Models as Markov Chains

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.812795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.812795Z digest=sha256:3d2911832865e23fdb9b2f437c8405d52071c28d1735d3a92733e90cd617288e

Observation c58d727f-2cd0-43b7-af55-33038c0db0f3 · outbound

This paper cites Star: bootstrapping reasoning with reasoning.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Star: bootstrapping reasoning with reasoning

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:41.160931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:37:40.815997Z digest=sha256:800f2ab2644406e5ac2ef9be64979b4f3a0f69b9f7a7121d6a7cc30eacbad9e9

Pith citing papers

Observation 8c284df7-dd23-469d-8630-b8d80b83a028 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

Reference 300

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:23.937668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:4c118b561d80e2193f265da68edffafd0512f81cb4eee7c256850de62bdb6b0e

Observation 2cd0d59c-5a84-4693-8361-82fbf94ab6ec · inbound

The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently cites this paper.

The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:24.330367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T05:27:11.761971Z digest=sha256:496bfcf494a419681a3a7c5f7d6d28a4767d03f975f4c6af79f00a866bda4882