Pith. sign in

Paper Citation Record · LEDGER

G-Zero: Self-Play for Open-Ended Generation from Zero Data

As of 10 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2605.09959.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09959 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:39:40.780801Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:47:26.709299Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.708592Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact28
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0bca19f-eea2-4e01-97e0-3a4d26c32514 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.388622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:3c6cca42ef1f5779596e4c8681de0b3be3a51b7b0eaed362798cff87af09781d

Observation 9fbd0733-d719-4076-b473-b0e906b14996 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:22.024625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:b18c1faf33f029f58d84333b42082be8c5901f0ba2da2f57f67dc3d967d3fd6a

Observation 31459089-a7f6-43eb-9146-0701db6d5c45 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:13:05.519270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:05b6c24d14a56ac398170d1fd74742e690859e689d3c2efd0cbe6c994c2d6c68

Observation c16415e0-110b-4953-8524-66b384c9247a · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Serl: Self-play reinforcement learning for large language models with limited data

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.384121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6d7a93d88c33f7372e8843562588da1b7920d5a21e887613374c61b811656d79

Observation 831c6896-fa2f-4a9c-89d4-62112c5abc83 · outbound

This paper cites The Llama 3 Herd of Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data The Llama 3 Herd of Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.398856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:3d5216da62373ec5fb8e82ca9f2abfe1e6ac5a728373a7552ff9e347f218ac86

Observation c0b57059-355f-42d3-afb8-5e498ca61b36 · outbound

This paper cites A survey on LLM-as-a-judge.The Innovation.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on LLM-as-a-judge.The Innovation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.687577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:941c72624edfa2dc94b44721238eafc30ae2c0b460af33d01cbe616df6f066d4

Observation 8d97dbb2-290a-45bd-ae91-22127eab1f0e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.407524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:1c75644e951ab171f6dc8ddc91e55c914038cd910c9dff8fb050eb7f3f4a0153

Observation 5f35d001-1f42-49c9-9a27-df899b525c8d · outbound

This paper cites Visplay: Self-evolving vision-language models from images.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Visplay: Self-evolving vision-language models from images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.411974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:3f2e0f626c4022b9a5edaea6dfa19e04dd8d7ae773ef18eca4edb4e3cfa3ce91

Observation 7422b71b-1af5-43aa-8217-53cd58313dae · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Lora: Low-rank adaptation of large language models.Iclr, 1(2):3

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.684032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:d6a339e386404b2f1e61f739d521fd8638281cbd795c41d0d01328a6999d34ad

Observation 00663b6f-c180-4836-998f-9ced3017c357 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:0cbb99d6b9529ff950f4bb2d4a773797499f0af3410e337785f9a28bd547e149

Observation b41f7614-9b3e-4d97-a544-1e7f5cd34ca3 · outbound

This paper cites Large language models can self-improve.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models can self-improve

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.675439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:c4ac3a9a0193d3993977c8d49ae0247a0a4870ea55c466096a45250fc95d40dd

Observation 57d026c8-67e1-4277-af20-7b32c25b4dfe · outbound

This paper cites Likelihood- based reward designs for general llm reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Likelihood- based reward designs for general llm reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.504310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:b4e947aa63a20b8efe300940eec22d21f1025b8d4442d3087389b12661e68bf9

Observation 28de626f-df92-4a20-975f-61ed2752f138 · outbound

This paper cites Mm-zero: Self-evolving multi-model vision language models from zero data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Mm-zero: Self-evolving multi-model vision language models from zero data

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.486410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:e7e8dba3f1ba0cdd1e971d9bf05068723f2c417ba2987538e0ab16cc0ea2f6d8

Observation 16caeec2-af6f-44dd-a223-a6cb3ee161ec · outbound

This paper cites SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.431684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:afddab73cd1034c65b5f14f7f1181a9e53b0d2a6db75397cafe34e6936e5a7a7

Observation 0d67e572-25c0-4647-95ca-7e215ba60fe2 · outbound

This paper cites Learning to solve and verify: A self-play framework for code and test generation.arXiv preprint arXiv:2502.14948.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Learning to solve and verify: A self-play framework for code and test generation.arXiv preprint arXiv:2502.14948

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.473026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:79e3c85989c3a5c0360cdcb52ede4cf5fd442024f857f7cdd091b495078351a4

Observation 87a2e22c-a01c-4e04-affe-7eec69fa1f5e · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Spice: Self-play in corpus environments improves reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.453062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:31be17390533831e0cefc255ef505e78739a1d4e1a7c456ea79e7607437b635a

Observation 1f0d9409-00c5-41f6-9c9b-641cfa2b5484 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.666766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:f60f2ab257291ab6cb8229524a28ddf5f62b582c732d9d2b2f01e78cd0af3f5d

Observation fbe44300-a37b-428e-8437-8a4214bfca15 · outbound

This paper cites Nover: Incentive training for language models via verifier-free reinforcement learning.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Nover: Incentive training for language models via verifier-free reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.670797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6d48acaea1c2ea332b3cb1481c9d02e8e57b1a328a78b7d7611d49b297cc398e

Observation bf0ff9b5-e832-49a3-9024-be7fed9b2648 · outbound

This paper cites Efficient paths and dense rewards: Probabilistic flow reasoning for large language models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Efficient paths and dense rewards: Probabilistic flow reasoning for large language models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.511550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:dc4593ada90e2761b56f2b9f3d396c1717529a793fe74bd2297675bb4533335d

Observation 83dfbca5-754e-4658-bc09-993207f89f4a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.553716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:b1a4a7227db458d3d420987631607aa94b50ce66ff07ca6a8de220bc92152e8d

Observation 8f05cf74-1770-4e58-ab9c-1663b97e08a5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.660181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6a8d1e733aeb05e20da7f040b39f171814b1d0576e5a3757ba72ecdd8c3ba0c5

Observation 5514be62-3bec-4cb7-9ecf-8b7dc3c10321 · outbound

This paper cites Can large reasoning models self-train?.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Can large reasoning models self-train?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.464065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:b5f845545168700351e52d6b1ba51757b816df4ff6f570e6d504419c3ab10a46

Observation 27103efc-20fa-49ba-9545-5f22c98f9388 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.437860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:c0a2f292d6cb5451ad4e0fbbcdbe5dd7afc4ac6dd22e77cde0cdb3cb0bee1ff0

Observation 8ca4eef0-e1bd-4d8f-92e3-b750f6589665 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:91b72ed0f21785f12b7ab0f91585c43cef39d710e67d39ab9497096fd858e8e0

Observation a998ee1d-6dcd-49e0-ada2-e292fee87927 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022): 755–759.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Ai models collapse when trained on recursively generated data.Nature, 631(8022): 755–759

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.680122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:ac0c373353ab2315f2ec36234ca1f2e44988de70c9ab56a1bb4088a5bf96c302

Observation 5a81e153-6fde-4b76-89e0-dd9a1e24c30f · outbound

This paper cites Large language models for data annotation and synthesis: A survey.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Large language models for data annotation and synthesis: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.655845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:188ec3284c3184c8f42e6d7a91d8c40417825def5da5fb57d341e6b09d310a2a

Observation 73f7c354-7309-4110-bfb0-f21c8baf3e9d · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A Survey on Self-Evolution of Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.492516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:ee07f16e7a488a2114f505b264f91f95093039d70cf042468ec0ede5432d9272

Observation cc76e99e-c79d-4848-a707-cde085ffddd0 · outbound

This paper cites Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.516462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:13664c84c7726ef4ecfe5a9eff58e3de67ddf5ef70a6fcde7ab562a52e5ad823

Observation 8494457b-3874-4abb-8e43-6112c20455ef · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.643799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:d06dd70deb75ad5a5f5e36d50d08eb169d49d70802ee26e7373886d0e0f462df

Observation 732e4e6d-91f2-4048-8869-e1e5ea745ace · outbound

This paper cites Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:24.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:32c9840cbbfd5c05e5a0bfc6d1f90e8a37b50d1c7569ccc1f252e321eb9906d2

Observation b2592111-9a7d-4043-969d-5886727b9b42 · outbound

This paper cites A systematic survey of self-evolving agents: From model-centric to environment-driven co-evolution.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A systematic survey of self-evolving agents: From model-centric to environment-driven co-evolution

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.647847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:79b99a909a576d0fd07566c6ef2c3bc42ebb7005245dd536612e93b78a6d3a7b

Observation 140089a8-92bd-4ec4-829c-5d4be0611c28 · outbound

This paper cites Reinforcement learning with conditional expectation reward.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcement learning with conditional expectation reward

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.534561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:0408f846864a45e74e703850a64d6c5adf08a1ff96d05ff4fa0bd13c8134fd50

Observation ee43180b-99cc-437e-894f-55660022b433 · outbound

This paper cites Qwen3 Technical Report.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Qwen3 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.477601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:672591d43d518077a2ece6864706d8010bb06bd0af9feff994fa793fa39ec27e

Observation 287766b0-65f8-4c3b-9523-f4e7edd5de08 · outbound

This paper cites A survey on recent advances in LLM-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1–38.

G-Zero: Self-Play for Open-Ended Generation from Zero Data A survey on recent advances in LLM-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1–38

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:41:47.651818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:5671654a4ea5414f658f57da1d697cda815ed24242078d5473f34e68acd58446

Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.558508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:d58ac618141592dd917211eb994298416448a4139ad697817b9b8de4610be519

Observation 23f2d637-7416-46f1-805d-7293ba346d2e · outbound

This paper cites Guided self-evolving llms with minimal human supervision.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Guided self-evolving llms with minimal human supervision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.457773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:75dd6835923f04f59b6932e7901f413055ebc16e5dbd208f9f28b9542a67e4bc

Observation cbea89e8-be55-4407-a99a-d11bd8775cdb · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:23:09.397455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:86a8afd18bb46a02f5faeb94d9e8e592a3feb688d931fbbe949ba0668638a43b

Observation 4b644172-508e-49e9-a4a0-88c20667348c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Instruction-Following Evaluation for Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.444360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:284dfda14f9a864a9dc9d0b2f44787ebdf1f0f95e7e3455b62d82091e0df6c15

Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.522825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:d6d271694831729340423300b1fe25a03bf65b90c5a469d3ac61c96e18a6e87a

Observation d6a93fdf-ad16-42ef-be51-5363fbcb7baa · outbound

This paper cites Self-Challenging Language Model Agents.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Self-Challenging Language Model Agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.540151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:2d32953443c9e2f966af8efd7e80e650247d8147f41795b6da72b8e9a3daa666

Pith citing papers

Observation 8e636395-6e99-48a7-bad2-df97c3671a52 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops G-Zero: Self-Play for Open-Ended Generation from Zero Data

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.709795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:2cb4f9cbaf5f65eb4918e2128545ffee5e4c5cde775c99e7621a7ac009f07b6d

Observation 110f500a-8414-4793-aa78-b431ad1af83f · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning G-Zero: Self-Play for Open-Ended Generation from Zero Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.493350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.493350Z digest=sha256:4cee09d238bcc73efebc22d1a14dabcf676579c884d3545bd35797e53aa75268

Observation 0622efa6-7c2a-4764-8ba0-0ff829eaac29 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning G-Zero: Self-Play for Open-Ended Generation from Zero Data

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:26.709299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:26.709299Z digest=sha256:03771ec21a2617c0523772ac728c53b7c663b49c6fa06def6f951c0f6172f02b