Pith. sign in

Paper Citation Record · LEDGER

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

As of 17 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 13 inbound Pith citation observations for arXiv:2504.19162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19162 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:14.799234Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.285830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:26:27.242958Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4818701-5414-400a-b2b0-3172ff58f326 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.505744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.505744Z digest=sha256:4d14ef2dd949dd13ffc5c4466d11c9258ed6c96e2a8107ced27f5ab0bf5676c7

Observation b5ce3b57-5804-4880-a878-40df3552ad78 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.510381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.510381Z digest=sha256:3e8cfe2963bb6c42f8df14bfb613a1f346b6f4cd2305bb482d9528e9be68ac1f

Observation df0345db-cf64-4aab-a358-bd463743294f · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.514628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.514628Z digest=sha256:e8f98a507c489282215e8670e7147f1221d7709f83ad7846b43784ae8ec98222

Observation 3a7747f8-1245-4caa-ae39-9760be751873 · outbound

This paper cites Language models are few-shot learners.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Language models are few-shot learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.685549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.518559Z digest=sha256:61af4f74b215bc679f911dd9048311a0bea1c30cedb34f76a98a7e5f94bb310c

Observation e5305829-e3d4-45f2-93a8-630535500090 · outbound

This paper cites GPT-4 Technical Report.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.522368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.522368Z digest=sha256:ebc1fa640de47ec3133b61408b7ab1aa757518e5f449c00e8d1459b8df7236b6

Observation 8336a4cf-a945-4d4b-9779-6e19f95b6cc9 · outbound

This paper cites GPT-4o System Card.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning GPT-4o System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.526276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.526276Z digest=sha256:7d59511dd3c821f7f2fa547c9301218b5ca547d0f1e4fe125fc0fac81b1f0f75

Observation 98957b60-8a51-4a93-9994-04fe5aab8d3a · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Gemini: A family of highly capable multimodal models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.673877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.530539Z digest=sha256:8d4d7efd9275219e8da7057e519241c9cd38345da359dce66a4ff2ae26910429

Observation 126b6d6c-fd87-4997-b62c-e67baad1b90d · outbound

This paper cites Introducing the next generation of claude, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Introducing the next generation of claude, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.534386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.534386Z digest=sha256:d7e60920513728bb58a886578749089d02755a99ae2459961b98882a06cf91aa

Observation 1b22498d-faa8-4b2b-9cb2-cd8533744021 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.538558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.538558Z digest=sha256:8a7d16516cdaa21d0a8dabc6f4ea69666a1b0a78249cf2105e17f8bee77e87ea

Observation 71a068a0-3664-4b5a-aefd-7e082174e3b3 · outbound

This paper cites The Llama 3 Herd of Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.543126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.543126Z digest=sha256:d1da9717ba43266d7be1d7e37de9b9d6e4ce39d03e7516366645eac611571ec8

Observation 3940a032-2ae3-4bf8-8e44-47a30610ba77 · outbound

This paper cites Qwen2 Technical Report.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.547292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.547292Z digest=sha256:8aab4670b79680f2d9b7764e8a07411f3cfab341cfa4dc80e73c10228a7143ad

Observation e6b8881a-ede2-457a-8d78-14d3f6f15d67 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2.5: A party of foundation models, September 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.551312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.551312Z digest=sha256:3a3e492b2b57cc35dd298af5f69d86f01dd6ae2d60bff149792a62b5ce071ce0

Observation 607bad48-0883-47b1-9c40-e25489c5fb88 · outbound

This paper cites DeepSeek-V3 Technical Report.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.554833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.554833Z digest=sha256:9241d961e2055e2d163a859c0a1e5ced92c246f62f9c5d1a755df2d246189730

Observation 3985ae24-e812-43ad-93f1-4933b7637891 · outbound

This paper cites Training language models to follow instructions with human feedback.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training language models to follow instructions with human feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.559247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.559247Z digest=sha256:38dc219cb5f24cb4b111b44585a526fdb288712bdaf2b5aba918d7965e0d4d20

Observation ce7406d9-8776-4bc6-8d1c-8bd8d96dfce0 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Scaling Instruction-Finetuned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.563002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.563002Z digest=sha256:1300c7b4044a981e87a05783116235c211ca86d73e6f54b4cdcdb4f2a0e04a0e

Observation de4e5c97-5d48-444f-a6d3-4b79199e8eb8 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Fine-Tuning Language Models from Human Preferences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.567228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.567228Z digest=sha256:3fafde0cf59f747bc8ec2889ac3be5e278b73fe29f1b6adafce41166195a9de8

Observation 3c82e5a0-b269-412f-9aee-d82439a078bf · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.571617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.571617Z digest=sha256:67b9491d95add361190e72f2f412398f630a194dfb5403714db33a08698beecf

Observation 81e9d7c7-abc3-4954-b1a5-30cf1d1eb324 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.575844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.575844Z digest=sha256:da9087d64b43fdb951ddde35b14f18b4685e061370b98e9d0219339cf7b241b1

Observation 659adce1-6566-4e18-9750-2302a2330d78 · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Prover-Verifier Games improve legibility of LLM outputs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.579837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.579837Z digest=sha256:fb2ceda9285ec63cf94eacb3bf22c82896d192118b76b60d6400dfe85d0eb23a

Observation 382acda4-e014-4b37-b9b4-75f5e90ccbda · outbound

This paper cites Openai o1 system card.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Openai o1 system card

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.641541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.584281Z digest=sha256:c7228f9b0dd6f8ce26043cb4547d7ddab31a6d3cf7d0ae2335b138b7d8fdcbfb

Observation 9ef91387-af08-46eb-acbf-d943021de027 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.587663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.587663Z digest=sha256:3b0cd3f289bf11b294c8bcd2bffb5b48e63d6f45da7d71ee60716eb3c8229ca2

Observation 123b766a-2c5b-449d-b8ad-84cc6dfb3316 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.591591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.591591Z digest=sha256:b260e854ece59d05a87612a6f5e7b0d8ed3b51fcdc81135a37ddf0a55e69e20e

Observation 91671b48-b76b-42f4-ba6c-5f99bb9d76e2 · outbound

This paper cites Let’s verify step by step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Let’s verify step by step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.595052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.595052Z digest=sha256:156f65e20f3236b0e5c907e022fc86d54acdeaf498dd76d57001715b61b35db8

Observation 14020d1c-0b0a-47ed-a7cc-a95d351da66d · outbound

This paper cites Skywork-o1 open series.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Skywork-o1 open series

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.615374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.598579Z digest=sha256:cf996112533b36b6fbd57ecdecd787385ad51932307669c5d6e72b7d394d3491

Observation e1e8737e-ea83-413f-99b7-d0ae2ae338c9 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback, 2022.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Solving math word problems with process- and outcome-based feedback, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.601970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.601970Z digest=sha256:1ab83141a6bc87cd10cef28bcb45af462eb81402c07e5686b8585d88d63a9355

Observation f6110111-5df8-4562-b836-aa4c29c75581 · outbound

This paper cites an unresolved cited work.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.605869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.605869Z digest=sha256:336e6849c9c33d07278916d3137100511f172410852fb84b4a65bf9d02d1d958

Observation 0e45cfc4-fcc0-49da-9398-c466e1de19c5 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.609396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.609396Z digest=sha256:15835d19409cceb7e6f32bbab44042daf81a42fb52d8054720c1e9c08b246737

Observation 64bbadf7-c265-4275-a40f-5cea1e0f3438 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Generative verifiers: Reward modeling as next-token prediction

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.588205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.613377Z digest=sha256:f259b237093b24d751b12f0824fdf6470d612f9df8a333287039c0f96cfff6e4

Observation e96f158f-e359-4e4f-bac7-52b63dee4c60 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.616880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.616880Z digest=sha256:b0312be867bf988a2dbe0dcfce5d96a2dc63ea9b39a665b27ae6772439e48384

Observation 6c250693-50eb-4c2d-9664-cc2b2f285829 · outbound

This paper cites Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.620811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.620811Z digest=sha256:9a4cadeaf3c577a45a05147951703b0f3c4ea953077675332484a04e21c82cee

Observation 27a71eaf-1c26-4cbc-b64f-d6301da3b8f1 · outbound

This paper cites S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.624752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.624752Z digest=sha256:350b91748af34cf00bac3ede19ccd66ad46bc390e2a42297da80be014fc41cda

Observation 2ee8533a-66dc-448b-8027-995dbdfc8a08 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.628433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.628433Z digest=sha256:4d02c43c17a8bd6813fec7bdd7111abe2fe3d547f284e585835148d3fd353565

Observation ac1f8437-b4dc-4094-9155-267366176814 · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.632334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.632334Z digest=sha256:948ea656ea0e164e5c88ff85e4cd9b8baf8fea5262340c207c9bc3d1b983c18b

Observation b7a22f54-3bf6-4827-9f33-52d5703701f3 · outbound

This paper cites O1 replication journey – part 2: Surpassing o1-preview through simple distillation big progress or bitter lesson? Github, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning O1 replication journey – part 2: Surpassing o1-preview through simple distillation big progress or bitter lesson? Github, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.576261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.636473Z digest=sha256:5e436f2e427b789fb6bcb3f34b6aaebf7d335d1df4b2ff45273e7c94839555f1

Observation dfb04b07-7747-45b5-955b-0e03d499233f · outbound

This paper cites Can large language models detect errors in long chain-of-thought reasoning?, 2025.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Can large language models detect errors in long chain-of-thought reasoning?, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.564637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.639953Z digest=sha256:5cc23c71a498451873fef99e01ea665b61fcd8b9ea27c262c09335b73aee4c95

Observation 7fc92d19-3714-444a-bfb4-c76381a12a9e · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.643907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.643907Z digest=sha256:c15d83f215401a047db9fdf301cf2e414bb8451702eaaf73d4458765cd3774ac

Observation 28f58e72-2a5d-441e-9b80-8c9f6b497f49 · outbound

This paper cites Aime 2024, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Aime 2024, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.551941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.647592Z digest=sha256:f9aaa5abf29572708f4cfb85783fb7ec3683995d891a64d1fa34ff7fe74dc084

Observation eae23f51-f64e-4364-94da-f84539ee27fc · outbound

This paper cites Introducing meta llama3: The most capable openly available llm to date, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Introducing meta llama3: The most capable openly available llm to date, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.540384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.650918Z digest=sha256:46c6a05a2f679478a1f71681f6435fdb57b2f4a45af1240cce35b8f2c4ef1392

Observation a0e24e55-c82a-446e-8918-7537347e26c7 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.654469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.654469Z digest=sha256:060589a7bb4db89ed452ee8a6a3f434cee3c207824bfe4c082d794386895f5bf

Observation 24849fa7-1d53-4019-b291-6823dd354168 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.657957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.657957Z digest=sha256:6fc1b2167064de8c9661698da3035e500540a634e5bd7cb2c10d01acd28f993d

Observation 6e1b4268-29d2-486c-b4fd-3ede57d6e91e · outbound

This paper cites Tinyzero.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Tinyzero

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.661555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.661555Z digest=sha256:791f32ca4a9f770cd6765fce6e669e9b48299396a10e17a8910a1dadba7e9a3a

Observation 8716aa30-5ffc-42c8-966b-7b000b68704b · outbound

This paper cites The lighthouse of language: Enhancing llm agents via critique-guided improvement.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning The lighthouse of language: Enhancing llm agents via critique-guided improvement

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.665289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.665289Z digest=sha256:4dcdfa5090352d4db9d2eaa3abec00d51b51172e4069bf467939ff0c7423331f

Observation 6c49bf63-b568-44fd-9b63-e71d0ac3af5b · outbound

This paper cites Some studies in machine learning using the game of checkers.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Some studies in machine learning using the game of checkers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.668679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.668679Z digest=sha256:10eff5f5b68bf78dec25d87b9288d1dc5e7794eb230688afed6efe987858fa9e

Observation 90787b21-6b19-4cd9-96d8-029faaa6ee36 · outbound

This paper cites Temporal difference learning and td-gammon.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Temporal difference learning and td-gammon

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.501633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.672523Z digest=sha256:3cb3901f464e38982964a01975be144059b1c4b92eb9b4a138e849990270b941

Observation 7ced5c1c-4dc2-4e17-8a53-fcf4c196b2b7 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.676179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.676179Z digest=sha256:f792bfed33efeef99bede403204b90df7227bb502afa555a1d89a73251f317d4

Observation def38d47-ab18-456d-8563-61ab31747f66 · outbound

This paper cites Mastering the game of go without human knowledge.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Mastering the game of go without human knowledge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.490754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.680019Z digest=sha256:defe23881eb3114b8bafbad8c4ed4da0c8284a7ba0f53b3e131d56d551e93ec1

Observation e49a1add-e3fc-4072-9728-f90a8444246c · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-playing adversarial language game enhances llm reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.479356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.683478Z digest=sha256:c2ee255771916a725fefd1094ba7605283abfe5944d8d1c8bb528842d417863b

Observation 587296c3-cfe3-49ca-8001-53afc87b5fdc · outbound

This paper cites Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.687140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.687140Z digest=sha256:2af962e95cca2548060cc57d3ae671d9a2e7100d9631cf3fd733ebd97c70e35c

Observation 3f7a9601-fa4a-477f-94ce-b72c0b3f0ad5 · outbound

This paper cites Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.691014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.691014Z digest=sha256:699f90efdbc818dc5964034778d3dfa73fe8fbbd366de27a57da17600f5e82d2

Observation 523096d3-8559-4589-8215-367a6c681402 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Play Preference Optimization for Language Model Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.695026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.695026Z digest=sha256:0879e0b9dbafbfe367a6b1b339fe359a05e800165154baa46092c4246e24f5e8

Observation 68a38343-2479-4304-9cb9-667ff43ea9f8 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.698960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.698960Z digest=sha256:77b6b9eb4bbffa965a5f04c11e68419ed52b8120b2ec34cd71599b1dc33814b8

Observation 363c2bc3-6819-426d-ad6e-5e376d9c5d98 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.702907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.702907Z digest=sha256:f5294889546912317e96650a7b1434ca9a4e8af8f8f0a10fef6a3c6a3e898280

Observation b017651f-4cc6-4ed3-bad1-c3d8be3080c5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.706747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.706747Z digest=sha256:52a1034ea0e6d45c309d66de1981f57963f6fbdd43c771fa83171dfe8bd3cc5f

Observation 47dbf16d-0a09-4396-9816-f17e01c1f0f4 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.710535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.710535Z digest=sha256:1e05ecc24356fa7af6c2f040bfd753681e672963825aba40438563ef7468acd2

Observation 6bc964d7-207f-4dbf-babb-3dc0fd28d14e · outbound

This paper cites Qwen2.5-math-7b, 2024.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Qwen2.5-math-7b, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.468330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.714348Z digest=sha256:fca816735346985f671065fbeca49ae8fafb504d36cd4900c5ac2a518fa852bf

Observation 1c590def-d60e-478d-9be4-2ce55403d821 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.717750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.717750Z digest=sha256:7c619a11dab37355a5bcd4e54aeabc47e0d39ba0ef40bcb3b64a854e31d36dcb

Observation fe913151-b115-4f67-a1a4-4f9ece4fe7df · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.721523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.721523Z digest=sha256:c92f0a65b1996fa1b5d706d0a4d85cf5c3a262e61041e98f09cb446d600583a6

Observation 8016ad59-934f-4767-9b6c-82862739b37b · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.725398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.725398Z digest=sha256:a08e52c817a4d58e8ec12b1bcc0494624dc2e30fa21300c4751a85f4b043a530

Observation f3b4ea95-52ea-4912-8ff1-901514054820 · outbound

This paper cites Note: For each question, you will be given a reference incorrect last step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Note: For each question, you will be given a reference incorrect last step

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.423442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.740981Z digest=sha256:3d324c4f732d67caa4e362c567aef7f280c316bfcdf6b8c51d76ef30532b7a32

Observation 6bd5e618-e267-4f57-aae1-f049e5c14b2b · outbound

This paper cites At the end of the response, output \\boxed{{Correct}} or \\boxed{{Incorrect}} to represent the correctness of theLast Step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning At the end of the response, output \\boxed{{Correct}} or \\boxed{{Incorrect}} to represent the correctness of theLast Step

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.390785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.751478Z digest=sha256:b1e92edb209f638257bb29273f474b38f70a71814b90b3e224c565c8504447d0

Observation 8538f9b0-462b-4ed8-83a6-e40eed9d9f8b · outbound

This paper cites If the Draft Critique includes this analysis, you can directly summarize from it.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning If the Draft Critique includes this analysis, you can directly summarize from it

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.379384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.755268Z digest=sha256:3231c36cb3e105672639fd6c4f1a50fdfaef496a40e2cf4c585d7664ef59e075

Observation 37bb3787-2a1a-40a7-88fc-da73680c846a · outbound

This paper cites You should write a new version of brief critique here.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning You should write a new version of brief critique here

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.368837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.758769Z digest=sha256:5bea7c6219693cfbc95c66aa44744de4e34319e73b385c9a6056b6ee619de66f

Observation 53131bf5-e9cd-4903-a610-04691b79be61 · outbound

This paper cites Please draw a conclusion about the correctness of the Last Step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Please draw a conclusion about the correctness of the Last Step

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.357381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.762121Z digest=sha256:07a078bc6c94e7b10ace3f85210f3f477e53924ab8847025ae578d8559d45f03

Observation 31d32af9-ab56-49a3-813d-f9cff886c7f9 · outbound

This paper cites the critique.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning the critique

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.346433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.765564Z digest=sha256:b86c8d676cb4b4e0439934b6809addd7abe47aa96ca6d9308c127e809811eb61

Observation bc56fce3-dcb1-4c2e-a818-b110ffd114db · outbound

This paper cites an unresolved cited work.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:04:15.334974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.768916Z digest=sha256:d30e563e17cfc42ecac58031a024baeb6b2fd2f62c412b6048751c5ea37f2b57

Observation bd740ae5-ee52-47df-9226-cf956d234b14 · outbound

This paper cites Instead, in the Critique, you start with an analysis of the Last Step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Instead, in the Critique, you start with an analysis of the Last Step

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.323484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.772248Z digest=sha256:c7ce9b8d24fc5bd3af8cbd46b59dc16c45db4899f01367da5245c0af841c2479

Observation 340cf9aa-4dc5-4041-bbd2-66e65078d28f · outbound

This paper cites Y ou only need to focus on the correctness of the Last Step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Y ou only need to focus on the correctness of the Last Step

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.312478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.775601Z digest=sha256:a0e2238b98c70721adbce6055ad1443f82161d3dc8e1f4bf75771ab13c683342

Observation abdcb106-5899-4b5b-99e2-b66576ea4bda · outbound

This paper cites Clearly explain the solving process in the last step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Clearly explain the solving process in the last step

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.457002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.778765Z digest=sha256:f49c2990ae328ede0f6113d205b108e9883dcfbdcea0dfd9ba71225ea89f7c10

Observation e1c3e9be-5bd2-4dd4-8aea-8b3ff0618aaa · outbound

This paper cites Predefined Error Types.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Predefined Error Types

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.445709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.782398Z digest=sha256:5278cfb528faa294b18c576e46d08f991e095477d72d246fa169e48cbb0d62eb

Observation 63eaa358-2ed3-4195-8007-0df80ecc3b84 · outbound

This paper cites an unresolved cited work.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:04:15.434535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.785615Z digest=sha256:4fd4f656d95dffcbdc113d23035e1a86f317d239aa736e069a78654041448d28

Observation c17ea9ab-5827-4b49-9985-6adc1f3e8fe3 · outbound

This paper cites an unresolved cited work.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:04:15.300833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.788959Z digest=sha256:9e791cda22bfad6ec9ff8144a2bb6e308f3c06554275c36ab7ec53732018eafb

Observation 950be891-8172-4be9-b83b-1a8be41985c5 · outbound

This paper cites an unresolved cited work.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:04:15.412224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.792447Z digest=sha256:14eed36060ef27b6442233e37a7a495fe1bbf9e2151d922c1366da9accc6e4ba

Observation c351eec6-d67b-4299-8ee4-b40abf1eabee · outbound

This paper cites You should write a brief critique here.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning You should write a brief critique here

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.401947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.795736Z digest=sha256:d614c9ce9e5cda38ef944db2550af3dfcdf939d387c88bf4cd9a3991b481f584

Observation e47bb238-85ec-44f9-9cce-35007ee680ad · outbound

This paper cites At the end of the response, output <Answer>Correct</Answer> or <Answer>Incorrect</Answer> to represent the correctness of the Last Step.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning At the end of the response, output <Answer>Correct</Answer> or <Answer>Incorrect</Answer> to represent the correctness of the Last Step

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:15.289595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:04:14.799234Z digest=sha256:fcd421c6d49c78cb99c5cd381510ad53ea93dc10b712f128dd4d31bada1bed39

Pith citing papers

Observation 8f069e31-eeec-4285-8eab-5a9ed739cc23 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.285830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.285830Z digest=sha256:07041accbfc4f2ba63432d1fca9ccd74e62d4f27fac57fb5b6bc71fadfed7fdf

Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.410838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.410838Z digest=sha256:583bf6117308098af804e9a71ecf81dc9c3b71c1a7d46aa28e116e9086d4239c

Observation a60935ee-dde6-4400-8c29-27d6730082b3 · inbound

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models cites this paper.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.520322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.520322Z digest=sha256:bc326f79c49434aff2c4dbb067186e0b56baba54e4bf2a1b664e6c2c0793d528

Observation 449b5695-ac43-4588-9625-8b28fde75fd8 · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:30.674164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:30.674164Z digest=sha256:2b7de5e339baa87433b0cde3ae2b1cdced07ec0e96ec95869dcc5c94c5b114d3

Observation 6a359f50-1ac9-471b-9300-3a38ec70366e · inbound

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards cites this paper.

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:26.869430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:26.869430Z digest=sha256:9c80da468dfee1d43f9ba9249601ff2b9168f0da0d998c343d550935310409c3

Observation adff3f02-7490-446b-9e66-2fd40970239c · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:14.992423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:c07753303e0484f78d9654c1ebea76ea3674db21c4a9c2431f3110cf79982acd

Observation c95307ca-c6ad-4781-96d8-a1e6fefae90c · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.750232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:a83517849028f7ccf4ef8018d110d3de2fc055c1f55eba1355c83aa655de7566

Observation b8b22f47-daa5-455e-80bb-bc43f4b79248 · inbound

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text cites this paper.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:18.201366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:43857849e8becf1ec192dfab5a9fbc77c4fd82e2867f6b1ed52cfe6165f8ebf9

Observation a0bca19f-eea2-4e01-97e0-3a4d26c32514 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.388622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:1eb643a874bb2b751d0f2bf43491288b6e1b9b73c95ca12016049f11fecd38b2

Observation 63701c8b-1f32-45b6-84b9-d8b8cb3173f4 · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:53:49.190678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:8e23e60ec4e41fea5d15f215523e77d0fdf1662fd6e3d6d6641ca8c4955fcb7d

Observation 72136545-3aff-48b4-82af-61d346e61e27 · inbound

Pseudo-Formalization for Automatic Proof Verification cites this paper.

Pseudo-Formalization for Automatic Proof Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:29:42.216570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T06:25:01.098420Z digest=sha256:c9a576fbee6a95cc1fe1b6974dcb774a233d74a83b7d7653a3b7537c17e4b9da

Observation 33f4a043-fb6c-4b9f-bbc3-3c9edc09d1b8 · inbound

Pseudo-Formalization for Automatic Proof Verification cites this paper.

Pseudo-Formalization for Automatic Proof Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.676445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T17:17:48.969969Z digest=sha256:580c1aee1c8ead045b20488ebc3755d81f60f932093039b6beb864071854b74a

Observation 84fa483e-fc08-4289-b1c0-d2a7ebe9195a · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:27.244722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:2f50b4eb9636863f217a317caa0e0f9cb0df31978865853c903647c902c075ed