Pith. sign in

Paper Citation Record · LEDGER

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

As of 12 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 14 inbound Pith citation observations for arXiv:2502.02508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02508 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:49.117837Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.038438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.429120Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1229900a-f24d-4041-8375-33a12896dd33 · outbound

This paper cites ( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.209342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.824996Z digest=sha256:3676a8c8118032302f37804f1a6279d0addf75ef2a46b7492dda10a4cb88b014

Observation ad843070-ac30-4804-a4dd-5e40acc01e59 · outbound

This paper cites Self- consistency improves chain of thought reasoning in lan- guage models,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Self- consistency improves chain of thought reasoning in lan- guage models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.754123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.754123Z digest=sha256:406f4f45886c16398107150932122d490e208f6f91d4a5c5d06833933a382773

Observation 22b49c8f-5b87-473f-b689-220da3b35a17 · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.758762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.758762Z digest=sha256:d8a706bf1d47510f28cc24637e64cbae46783d0a1ad8d703ba1c1233371c3087

Observation e4c20d2e-04df-4ed4-ad19-030622c80607 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.763506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.763506Z digest=sha256:7b17d89f2ed713c7ae92c6c058f5f285574fa1a4624ec7b0c086e4dfecefa7c9

Observation f07620ed-7ca1-4ac4-9df6-786b49367384 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.773155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.773155Z digest=sha256:fd52edbc5faab365e95dfee6e36c8272017e6e7c63ae41529d9ba2039c148fba

Observation d2ac63d0-277d-4136-855b-3533293a6cef · outbound

This paper cites Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.778636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.778636Z digest=sha256:bde78611346d83f93c89066991dfded91a79bdc2a0b2d93f92e0150510c36bb1

Observation c692f436-95e2-4d79-8f9f-7afaea91f86f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.066840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.873047Z digest=sha256:0eacc497b8b1fc790106d58cc41189118fbfe63f2d80b4242cedb2a47349dfa5

Observation c1d70402-b030-47bc-b71a-db2d3020a1d0 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.053463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.877128Z digest=sha256:931da49f685ea5754988642aa3325d461b461f0eadc27277963e19bf5df2705b

Observation a4e8493c-1a30-45a9-b0bd-7bf9cab9f7c4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.040040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.881086Z digest=sha256:41f99f77648fc81c6e0ce1477fe2a00f867c2dd85ac5f0c1ea8411002507f0a0

Observation 7f7acff6-6204-4b09-a876-a4dc6805cb13 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.797404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.797404Z digest=sha256:b50ecc495ce7d3160f095fc5e9f5ef95bf82b23725a0779dee110ec926af8332

Observation 9d6d1d30-2b9b-442a-8479-e85dc2cdf8fd · outbound

This paper cites FIMO: A Challenge Formal Dataset for Automated Theorem Proving.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search FIMO: A Challenge Formal Dataset for Automated Theorem Proving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.802192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.802192Z digest=sha256:9d9299c5ab7e6d6666ef04e11dadef7fa3b61138c77cf80840c88ce89f5ec4b6

Observation cf00df94-b20b-4b68-8a09-efa45ca0b981 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.815737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.815737Z digest=sha256:612f1cb5ad8c5e51173436a259c70af52bced10ac03b62f42af6c89f63b3c79d

Observation 17b21789-6b6d-49df-8266-65ab88b727f1 · outbound

This paper cites The final answer is: 7 Figure 6: Math Domain Example.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The final answer is: 7 Figure 6: Math Domain Example

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.820421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.820421Z digest=sha256:154734c71bbf99744e30b61dec71fff0aa6ff23dbcb1594c02dc71428cd8a524

Observation 4951b2ab-7486-4d87-abd6-a331eddc6b91 · outbound

This paper cites - For p = 2: The exponent in 17! is 15.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - For p = 2: The exponent in 17! is 15

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.170634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.838795Z digest=sha256:37c39b5d8964270e997b0a44940898975fbbd98d336e35731de6ebfcb84e0896

Observation 90b6fbf8-c42e-458b-be63-406341098dd8 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.196434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.829328Z digest=sha256:bb5444f5d0a416742a1ea13ffb8bb427c6e7c1f70935c684acc4732ad24bab5d

Observation e10216c5-ea86-4ad2-832d-5319d6692a62 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.157546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.843189Z digest=sha256:b5dc7fcb59bcc23c0c5ad9cb4e2f52837944ae2fdb2c82ab1fd57985aeee28fb

Observation f2b2bf08-6002-4209-9eba-b358a6efedc4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.144768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.847508Z digest=sha256:8d3f9993afc305d40a495e025fa4a6cbe64cc40432d81cc3cc4bb271ac17664e

Observation e92b71bf-6983-4a18-854a-7f2ad207b046 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.131974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.851794Z digest=sha256:2bb1c5bda489633cd8ed7ffcdc2c334a93ecb0cc329aebb2f8d1035aaab48856

Observation abb21582-f2a0-4546-bc1c-a6fcfe9e8b39 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.119153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.856150Z digest=sha256:675e4032a4dcfe93f6c3aefdd780cc9d549a570965615b32ade60c4a64c63457

Observation 777bdcce-ada4-49fc-97fb-1b9de9789362 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.106481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.860446Z digest=sha256:7abe20f7ddc88c8fa21d8f54857d4e7f32bad024377acb5dbc6701830d61e3e3

Observation a8672318-e70f-4ff6-8aea-480f3952d479 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.093349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.864677Z digest=sha256:cd12bc1bf8a2b6daf8af46c62326643d3dea2b0da5462a4c2673eb5902a76a9f

Observation 87438d9e-6158-4009-9c43-d64daa440b31 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:50.080011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.868939Z digest=sha256:3d6abce0dda1cccb327eb148e3fd82f25633a73b73cb215bb1da0db7e7a826ac

Observation ce5e6b8e-8bcf-492d-850e-299148b484f4 · outbound

This paper cites - The liger is a physiotherapist.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The liger is a physiotherapist

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.027233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.885162Z digest=sha256:25f2718143e223520748254bb75eece33806a3cf9bcbd151d64936bf14db8c84

Observation 02121093-8594-4fab-969a-83c14b37d159 · outbound

This paper cites - The dimensions of the box are 52.3 x 43.6 x 36.1 inches.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The dimensions of the box are 52.3 x 43.6 x 36.1 inches

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.013742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.889206Z digest=sha256:0e083c1bd0b5a18ee22d65d08d6cd28115fd916477cf63ab203f5de5a09d1ad8

Observation 9523bf24-98d7-47d6-bb58-8f9ecaa43a15 · outbound

This paper cites - The seal hides the cards that she has from the bee but does not build a power plant near the green fields of the husky.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search - The seal hides the cards that she has from the bee but does not build a power plant near the green fields of the husky

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.000605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.893289Z digest=sha256:e3c8395a86c77eff40d04396a95360e30ddfa355e14239cd7f4601418b7139a0

Observation bf68543f-2169-4be7-8dd4-c5c382ab41d0 · outbound

This paper cites Therefore, the final answer is: True.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: True

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.987513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.897435Z digest=sha256:113fa5c4bf3bb64833b46878db3a1cc4209ce7182e447f1b0d3621e48ad483c4

Observation 44c8e436-d893-49ea-870e-02b0eb7a74ef · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.974480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.901882Z digest=sha256:190ef0579563ffe9b3bb1edc88fc2e0de26b407719c7341f0159299b5d85e212

Observation 3571d180-50a5-4a2b-9753-e55d5e86d7e4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.961712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.906295Z digest=sha256:a8e43574560ef255d1433c0f03409bcb89c65e57bb354d85ea0bf4a9afcaefee

Observation ea7231ba-b934-4bd9-aba9-6f1de521ebc8 · outbound

This paper cites Based on the facts above, answer the following question.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Based on the facts above, answer the following question

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.948650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.910561Z digest=sha256:797a5ca90e02a1bb467fefed86ccab0bf0c0de0f162ef7867f6c0f921240187e

Observation 99591b1a-de19-4422-93e4-dbfbe821be49 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.935352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.914960Z digest=sha256:e17670d9f9b36c58f11f0118fb33f98a809701549ee23cfc35dae3c5243f4149

Observation d4f71e77-8491-4029-ad36-21663d827ae3 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.922571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.919175Z digest=sha256:df1a9e7db72ca09cea26d294068be7f6efa774f27b42b321723e65247620e5f0

Observation 9c233a02-91d4-4681-a439-8cff21c8f94a · outbound

This paper cites Christopher Reeve’s spinal cord injury was severe, and he required specialized medical equipment and ongoing treatment.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Christopher Reeve’s spinal cord injury was severe, and he required specialized medical equipment and ongoing treatment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.909191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.923486Z digest=sha256:266613dfc7571983a20bf049c2c8b7d67318d7b302f51e52d148d4edae547300

Observation fd1a1377-5d53-4f12-9b72-8680a4a240c4 · outbound

This paper cites The molar mass of Mg is 24.31 g/mol.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The molar mass of Mg is 24.31 g/mol

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.895865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.927807Z digest=sha256:582e51bf4dfa2b4a7b69eff2cc7864c19cc2053e968ab663c6b98037bb5d0828

Observation 0dcf4edd-ccde-4e4f-907c-e2be8e9c3cae · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.882898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.931959Z digest=sha256:1903c0e65bd8b69426809467ec6bfdf072ec296aaed70b05795af6ca13014344

Observation 69830b52-fb84-451f-88ba-3d86a1614d41 · outbound

This paper cites Since the reaction occurs in a beaker and the volume change is significant, we need to consider the external pressure and the change in volume.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Since the reaction occurs in a beaker and the volume change is significant, we need to consider the external pressure and the change in volume

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.869824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.936376Z digest=sha256:e52cf76cd3fa19694b2a404377af9556f854b8e3c7c7142a95520f3815cbbafc

Observation 2c19af94-8fab-42ae-8de9-effd2345dfb4 · outbound

This paper cites Here, text = ’ertubwi’ , sep = ’p’ , and maxsplit = 5.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Here, text = ’ertubwi’ , sep = ’p’ , and maxsplit = 5

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.856599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.940664Z digest=sha256:215a85587356a70f31a08b170cf96e7e9853fad0a02213ae67c58478a231035b

Observation ba60185a-c98c-40c9-a26a-1ba0eee35b54 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.843271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.945072Z digest=sha256:03efed68a26a9005e1d1e9bb351ddd6f8a1cf603f4b763e4cec807d9cf31cce2

Observation ec314509-d6d9-4609-9331-5d69742a99f6 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.830571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.949270Z digest=sha256:b691ac80c0380cce3ead44c1eb0d324b450bd4e2e81bfcb77f65557a5b31c48a

Observation b6ca5d4f-7b25-4fdd-856f-60c266f740d1 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.817832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.953461Z digest=sha256:f3e9f9b1fc02579b0a0a08020f61deb96dfa625bbd38202296effc0daba34e2c

Observation 349d8217-d6a6-45f3-870e-cd1e5e910ed8 · outbound

This paper cites Let’s consider the correct approach:.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Let’s consider the correct approach:

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.805212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.957469Z digest=sha256:0c1d1b067afb328d8f0c0cbedc23376fafd1e967598e88ccaff38fc62ce7e44d

Observation f816ac15-1d9a-4c7d-a835-3929024227bd · outbound

This paper cites Given the function’s behavior and the input, the correct approach is to split the string into two equal parts and reverse the first part.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Given the function’s behavior and the input, the correct approach is to split the string into two equal parts and reverse the first part

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.792615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.961799Z digest=sha256:63fe17dbc2f90ddf337777ac48aafdd98e07c29361555a036993e874b4c46f3b

Observation 86edcf3a-a8eb-42be-8f1e-bd830876c99a · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.779702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.966290Z digest=sha256:e5c7ce6007d35a893d74f784499f1643ae660087c9ab97c701cba49854c0be61

Observation 8ecbffa9-686f-4070-b627-d7e3800d4444 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.766859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.970794Z digest=sha256:a480cf32aa5d84da2f54207d67d79056fb4919a6f460560fb1bd5ca7fbec7111

Observation 05b90b56-5fab-4fb8-aca9-12731434e318 · outbound

This paper cites Therefore, the final answer is: uertpbwi.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: uertpbwi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.754569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.974976Z digest=sha256:8bb1e40319b7d35f1329b8a474e10e499b2fc7db77948bd4d495e5eb660895ab

Observation 1ca2df08-b4d2-4b80-aee0-0bf7aad63402 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.741549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.979367Z digest=sha256:60057b19b1afb34e319cb346e878943e14cc52e7d095091241496cc18cef3ae5

Observation c5f63040-891e-405e-b8bd-d9d959bff08f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.729279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.983809Z digest=sha256:0570f3246d76c0958d19c41ab6b8f44d1fdb6a428d701c1a328e2c98ccf63f81

Observation abb1883f-114f-4317-9126-d9c4c5187dea · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.716694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.988466Z digest=sha256:562f90769a0a2be0cd707e33821bb6ba48aa96514a320d3b78f86562bba9c0ee

Observation 3e88025a-aafd-41ca-a416-84d63f7d2690 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.703827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.992777Z digest=sha256:f055c5902e0671b1aebcaaaf7cd7ce2f14206119ab0ed840efbdf0a0c5a53b89

Observation ceeaa65e-5ae9-4037-8c03-29b00904140d · outbound

This paper cites superhuman.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search superhuman

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.690490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.996928Z digest=sha256:e4f79093d769aa6868200702d8c051af5e0eee5bfdc99c5fd46f2cb1ac0685b9

Observation 825bd4f2-9bbf-4372-8d44-9d73c49354da · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.677070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.001527Z digest=sha256:0bea61c240828e133603614039e8fa8d2f9d0d1502235c1771fa8e0fbb9a3836

Observation f374d70c-f487-4960-b004-b9010dd17816 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.663891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.006055Z digest=sha256:9937afdfb04bc477352efaf10a8b19426d1eac6bb5e0002dc98c0778ce4cc1ae

Observation 440fd80b-05ef-4887-94d2-c5c0950a5070 · outbound

This paper cites $x$\", (high, 0), E); label(\.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search $x$\", (high, 0), E); label(\

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:50.183460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:48.833684Z digest=sha256:7822a44bedc040aba3eb1296a7161dc8d7deef285535fd973989c22eaadb93b8

Observation a98f65f2-ce8b-4bd4-9cd3-4a3eb0c65d48 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.650716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.010468Z digest=sha256:b7d783d8d87b22c49e99a8fe2857b7fb225f22f61e61c87008fd5d02bddd67aa

Observation 6acb6765-8bdd-49c4-8d36-86741e19e657 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.637406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.014684Z digest=sha256:585007bdc556ce207c758d177987cacf9d6c166d92e685d4de4fc13c101419c3

Observation 8cc3729f-f38e-4e13-af50-12adacdcab2c · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.624675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.019125Z digest=sha256:1c920f6cb4f1ac211e991da874e8c71c62b5b95cc196337821551a583f368c7d

Observation fdf46ab8-92e1-40e4-87dd-0dd07027af26 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.612216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.023421Z digest=sha256:03ae8215d920f706b8d7f3dc20c3641ee0bd1ee33cd45bcfc1c8513abe3fe678

Observation 3bb8c525-c608-422b-bad7-b8df7e631eb8 · outbound

This paper cites The prompt templates for these situations are detailed in Appendix D.1.1.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The prompt templates for these situations are detailed in Appendix D.1.1

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.599320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.027700Z digest=sha256:316a07e26119161a3f9b449a9d03a559e89bf5556fa73e31912a6444bc02e598

Observation 73504001-38ec-4431-a69c-b15472a3f8fd · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.586401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.031908Z digest=sha256:66852194098d15b4585b3093d619fcc088c10ad2d9259e82f3237fb6d6b88ea7

Observation f9036570-0158-46fc-a653-4793b71f957f · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.573534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.036034Z digest=sha256:8719068bda9d10fdeacf73464f0103a7704de606fed399747528d2cadf4794a4

Observation 2be8cf18-cc91-498b-8c66-97132679f4f4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.560445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.040666Z digest=sha256:3de6e173439d9dab45f2f57f3cfaa75dde1148a07bf23e5de85f255eb375d043

Observation a31b254e-eaf5-4416-b473-50d75e761889 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.547163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.044888Z digest=sha256:07d02ce27e2ee4b7f36539bf09102cb4423f267a4ff132ba7efb6803098a78d6

Observation 21746f7e-f1ff-437a-9a0f-a5188a9b3e90 · outbound

This paper cites Therefore, the final answer is: \(\boxed{answer}\).

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Therefore, the final answer is: \(\boxed{answer}\)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.533673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.048976Z digest=sha256:0a933e5a70c89029f1938c13084e0aa2baa28a85ce63b3b5171caea49e811347

Observation 74f75386-9f21-4dbb-a897-5a9759c8299b · outbound

This paper cites Verify: [brief explanation of why you are correct with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Verify: [brief explanation of why you are correct with one sentence]

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.520413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.053555Z digest=sha256:fdaf23bf427c4a9e1a225e1dae3d862f22e8c7a950f7d006846a088984d19168

Observation b8140fec-c184-44ac-ba4f-ba4d59b635a4 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.506941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.057793Z digest=sha256:fd7fcb7e5d051d6c8ac11f2ba185ef9caae44bec456a74173ab8ae2069275798

Observation 8e954bae-307e-4996-a927-807d6d650c62 · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.492947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.061965Z digest=sha256:33cbf92410b17c0f1830aa2149550e69e3b93ee2dcab265f080c0a96bbb3566e

Observation 5145fcb8-f4d6-446b-b393-24040ba1a788 · outbound

This paper cites Your task is to carefully review your own solution to a math problem, and adhere to the following guidelines:.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your task is to carefully review your own solution to a math problem, and adhere to the following guidelines:

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.479751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.066144Z digest=sha256:56b0312476a3772d7ae69bc8226c5f857a260cfe096cbb5241056f4fb99f9ef5

Observation 001184a8-ab1d-465f-bf9a-98620a7691d4 · outbound

This paper cites In Step <id>: [brief explanation of the mistake with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search In Step <id>: [brief explanation of the mistake with one sentence]

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.466269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.070386Z digest=sha256:c84d70d893fbdcdd7099a6a0e5c5727ccb72502c5cdbfa5c5fa8f8c531b8be5b

Observation 175599de-f265-4d54-94bb-a9a071bf276b · outbound

This paper cites Alternatively: [your suggested step with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.452554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.074909Z digest=sha256:276797c6a5a0ac30c0a73c6c561fb3b12eb3b53009303bd81a48fa22eb4976f0

Observation 5ec0956c-14ff-4e15-9f1b-8914bc491e99 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.438562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.078980Z digest=sha256:fcaddb0c5b63aaec7c6b51722e0da020d2a584bd58b2c82c63cb10b7292d8511

Observation 12b31eee-316f-4c46-a17b-ac1890d98f6e · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.425042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.083128Z digest=sha256:487961af1355e0f8e432a2cd282256f3a0cff4cf61f62e02b4e6292fc699b588

Observation 87e3eeb2-83dd-4dff-89fa-3a4ca5684bce · outbound

This paper cites You are collaborating with a partner to solve math problems.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search You are collaborating with a partner to solve math problems

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.411650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.087244Z digest=sha256:385994d5110fd0aeec67647729e2ad3a8e045f6ccaace58a89a5a64358914435

Observation d0f29297-4655-444c-9f67-15c26dcb8c94 · outbound

This paper cites Your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.398107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.092476Z digest=sha256:9b4fbddd2635412337f87ac05aa754499e7b978aa8a80fe9465de46b12f48586

Observation 68dcfcd3-ac7a-4ee1-bcfc-b54d0dd5352a · outbound

This paper cites Your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Your partner’s partial solution

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.383765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.096709Z digest=sha256:04883ac7331d47e2a781d7c8c78cbf78674105d2e3a2aaf0d45c9e69582de7ab

Observation 49e2bf81-44b6-4ce0-853b-db0cf2e0b213 · outbound

This paper cites Alternatively: [your suggested step with one sentence].

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Alternatively: [your suggested step with one sentence]

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.368878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.100902Z digest=sha256:8cf41aa224fd81b4f910470b4373a576bd3d008fd5b00c2461ed53e07514388e

Observation 1eb586c8-faf3-4076-939e-fca111bc0d06 · outbound

This paper cites an unresolved cited work.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:49.354962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.105098Z digest=sha256:ffdbdc43b4455c01c0b71cd5ac252cab7283b345e6e36ce03be70d07ef80c9dd

Observation b8929a89-83dd-49ab-9288-3c0f08af1672 · outbound

This paper cites ground truth solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search ground truth solution

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.340855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.109751Z digest=sha256:e2f1c745b64062442e304b8bf5be87246632258fb78f07581102cd7d90239134

Observation 343bcb62-c7e8-47c8-acc6-b67a5f3a413d · outbound

This paper cites DO NOT refer to any mistake in your partner’s partial solution.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search DO NOT refer to any mistake in your partner’s partial solution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.327010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.113775Z digest=sha256:f71f195fa251efb74854c67f7c181d0bea78200706e967ec336ee85396de7c6b

Observation f3f5fad2-c131-4f0f-be50-f619e458dcd5 · outbound

This paper cites three two five.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search three two five

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:49.311495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:49.117837Z digest=sha256:1d4640adc49b21ede99fe82cebcd5cf9a533615c1e11874416168849be64277f

Observation 2bd64f1d-4468-408b-a609-8a5c9002de74 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Proximal Policy Optimization Algorithms

Reference 668

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.792502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.792502Z digest=sha256:a0695a4895dfd15b1241a04ef5d06d2701ebe2041c823d97c39a69ff12bef2dc

Observation 48295069-fb9a-405b-bdbc-fd320cc7f24f · outbound

This paper cites Efficient reductions for imitation learning,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Efficient reductions for imitation learning,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.788137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.788137Z digest=sha256:06dbc94e490edbd1ba139d56fc986910db633753bc818af67dac7e893ac1e86c

Observation 25f66ac8-423b-486c-9a92-82531b4a572c · outbound

This paper cites The CLRS-Text Algorithmic Reasoning Language Benchmark.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search The CLRS-Text Algorithmic Reasoning Language Benchmark

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.811186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.811186Z digest=sha256:bbaccd9cf13ce0ceba4bebc1598579e859770806a8f0c32c73aabe05e6641300

Observation 9b043971-f640-4552-aee1-adc2149f58f7 · outbound

This paper cites Im- itation learning: A survey of learning methods,.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Im- itation learning: A survey of learning methods,

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.783650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.783650Z digest=sha256:b640f7b37e7bdc03b7a7a583d6556abd1bf8ce105ea4e2c79f3cc30171b11a55

Observation 74fa24fa-07e2-48f0-99c3-4e3cb1325270 · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.748398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.748398Z digest=sha256:26522725a9ae23a774f5fb2bb4c5e88f1116743cdf872c107a6c4a1879c6e07c

Observation 18646795-6beb-43db-b12b-d781a9d6c42c · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 3634

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.806698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.806698Z digest=sha256:25b8d16780357bf5b139b1b711456d160ff76b56eeccb8ade1648dc8ba7592da

Pith citing papers

Observation 3dce8e4f-0f9d-4d11-a579-5c0aabe028d2 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.038438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.038438Z digest=sha256:085b1ab8b1a47e170ee6855076985645a49d9363c09196638372af9a21e3799d

Observation f9a615c3-b6df-4fb9-b809-2dc43c1778f4 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 247

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.423192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:da431f259e57ac581c0980ca31fc869af4f0d882a876e3c5efcde61d97bee03e

Observation 44ac4dce-7654-4d1a-88b3-5134cc8e77f0 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.324930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.324930Z digest=sha256:5c825b6423b663fd3cc5bad8d0ba2886906ac3ceda1b9ea72bfdcc6d2c8b9962

Observation e4ecac6f-8fb3-4663-a6a9-f731e82367d8 · inbound

Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering cites this paper.

Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:22.078456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:22.078456Z digest=sha256:9318a315de120090e1bdf2d685e6a381f3c7ac79489d977c49ad53439a7aff2d

Observation 9fce8929-7c3b-4dd7-9c7f-21a23129442c · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.425928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.425928Z digest=sha256:7a7e9487ece942cbc6023a1dbd41192e2e897ea98a7ffd535ca7a41ba348a623

Observation f604a5e2-42d6-4142-893b-39539c96e298 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:39.674751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:39.674751Z digest=sha256:9d75cd3e61046ee4d1a3156621ddb0f05332164eef9c0538ff7cc6ff7440377b

Observation d3877bd5-b1fc-4b05-8349-1711b53a8446 · inbound

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL cites this paper.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.395257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.395257Z digest=sha256:0f8078f60915ea07af5f3fc0d25e1180909961c3b6c830edfb89f618127973ae

Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.211988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.211988Z digest=sha256:5b0297993bd357c2de788929952d5528b3e0dae2b0e700da931657ffdc83a082

Observation 6c5bfa9d-980d-4cbb-980e-0309ac158e4e · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.932850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:0e6d117b59c90efa3faa31924002fd52544ebf84cee865c62e8c6e7a3e7d034e

Observation 95abce7f-b9cf-4393-a279-22f842fa3615 · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:49:00.696829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:05410a8b03b95010de3b18f5c70e44a60a0601a9cab07392c61a608f90111aaf

Observation fb1c88c3-4325-4065-858e-d40580c05100 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.990511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:cc8ebc84d6ed80656392f72dde2c83524dbd7b4eb0527f8e8ffaded487e85a94

Observation 4dae21f0-5cb3-4497-ac38-039c04402437 · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:41.015581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:20da4473d14652997f455b67f23193c0ea7cba0809ecca24d0db068f2b671652

Observation d65c1290-9258-4ad1-accd-ea2af3017a66 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.890131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:38b1925a03698e41e0657f4684e3d19ce8d5034754fcb8633e849860433765f0

Observation 545b8966-bd2a-4b23-b2e0-185011057ac0 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.430647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:fea04e8db687da3643bcd0ed760bd8a2df57298c5a9d1ae695269ac29bc41e27