Pith. sign in

Paper Citation Record · LEDGER

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

As of 4 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 2 inbound Pith citation observations for arXiv:2604.20051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20051 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:51:00.913166Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:47:18.868835Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact23
  • verified fuzzy41
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c16066c7-4b69-46b8-80e6-d3f021fd5220 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:18:21.444858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ffdfb24cce59a73141baac3e91454139583ccdf146b1a966e8e5e9b75d12de06

Observation b8b22f47-daa5-455e-80bb-bc43f4b79248 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:18.201366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:151fad663da4a81799f21f32e00e9c33fe8e047cc4b637190b209b836d114061

Observation 9b305eee-49a1-4e18-b7ad-5f8008fd23ad · outbound

This paper cites Self-playing Adversarial Language Game Enhances LLM Reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-playing Adversarial Language Game Enhances LLM Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.065828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:18f230c82d044cf99e2029b6dfdecbad30d403811d02202a91e94ce4c507997a

Observation cad5ea11-a550-4ee0-88cf-28b1382f298d · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Scaling Instruction-Finetuned Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:23:37.463337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:6a236a9048c0de9147a1058eb15cee3487340a92b2eabac619288e0d4bfe7e03

Observation 5ed87529-ca00-4062-963b-8737ce276527 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.791554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:f0f8bfdfcb9a4d2f6e8725b051c9748cd8416558584fa12ad315458e61583bcb

Observation ee773845-5881-47a3-aa5b-3c7013d6cb71 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.865950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:930568d20244658d6f92ea71612322654e69197c738578023ee80f748c87d13b

Observation 1b4f9982-637e-4fc4-b90c-afdd0c0ea026 · outbound

This paper cites Qa-lign: Aligning llms through constitutionally decomposed qa.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Qa-lign: Aligning llms through constitutionally decomposed qa

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.332412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:087f4e73239c7bb280623bfa43a735c5e7c8e91c3be645c5a68342749fb0548c

Observation 13ee7d65-ed37-46ce-85f8-ba1621619c42 · outbound

This paper cites Openwebtext corpus.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Openwebtext corpus

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.335087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:362eea90dda33f84a4b90434e958d0402000c9f379cc9145f590f7f73fefe8d7

Observation 4b045794-9b96-4880-9d6f-2330be434a01 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.884714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e82293aa58f207f5f9c92e7d8b339ca7ec30b8dc2844aef3bb8f02599ecabe38

Observation a50147df-d2eb-4762-bf22-a40cc296e96f · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Lighteval: A lightweight framework for llm evaluation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.356439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c5cb9fd2add8cbb4295e73ce616a3ce090c4a965ded1eabe68d5c83d611914dc

Observation f6e52eaa-d0ce-4952-b903-f137865b1e89 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text arXiv preprint arXiv:2511.10507 , year=

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.432193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:0487e0118a86e162460105f63ff8bcf9768d00e3f2d3509f25afe19e7ffec4d3

Observation ead073ee-97ef-44fa-9525-ef81d57423d6 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Measuring mathematical problem solving with the math dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.320248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:838e7c4ac081058f957589c547960b762406c508c11c4f9bcd8040930a14ed56

Observation 16ef9c3e-12c6-4542-bc13-dc746c5cc050 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:74bfc81401ca2caf41dbe5cbe5594300d070fb34b6e53ec9c292ee926c713fac

Observation 35977edc-4327-404e-9af5-ab41d386ea9c · outbound

This paper cites Dcrm: A heuristic to measure response pair quality in preference optimization.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Dcrm: A heuristic to measure response pair quality in preference optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.314439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:511d36fab1d39cdcd962f1f0e6c0b16798572a9d4149db4c99282ab5d7c73850

Observation e6e2648b-8406-4f60-a171-efc2ff33a10e · outbound

This paper cites Large language models can self-improve.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Large language models can self-improve

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.317630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:7d27e82873cc6296f5ad14dca9aa686f416d87d2f18eeda4a3a514a59c29f21d

Observation df2ee1d2-1693-4e1f-85b2-8b40ec090896 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Reinforcement Learning with Rubric Anchors

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.766370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c28b6f7790cc5c232f912758dd924673003f7999ba3a28833671b8f3bdff8303

Observation 46953652-3dbc-4c3c-b4ce-d0cec873f233 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:00:28.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:85d5f89602cbe4295dc1445ba3b31f6ae9adfbaaa1719ae4fc643cb317614677

Observation a918c201-3f10-4d45-90eb-574ba0ab1a87 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.322811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:2af4cace2e65769e7742e65ceaf66a3323a0980edbb09e1dee12ad7ff711deeb

Observation 13561926-7329-4bec-a0c7-7d0a9ae73aef · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Truthfulqa: Measuring how models mimic human falsehoods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.306823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:90c0e3a2d06da7c1c157bf54a422485ce4c9837f09407cfe243b7f1893eeb720

Observation 44782e97-3381-4c8d-bcce-60eb2d38ca20 · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Spice: Self-play in corpus environments improves reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.426077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:2265a6b17f98fc967279fb21fa04facaa69b8056962a31fa36a59b88e8223a9d

Observation 0cebd866-f376-4158-9a07-dd5a5d0450fd · outbound

This paper cites Decoupled weight decay regularization.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Decoupled weight decay regularization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.297173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ff25df1a896ac5a687bc6ac3bcdc85055f125775b2abbe5f7b7cecdc0405b17f

Observation 23004da3-0573-4573-80d6-05c6ebcc779f · outbound

This paper cites Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.354128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:93b1b313a92102628b275f5740fb9a69d79488cef4e9695d655ab789cec30c0d

Observation 6cc02ad7-93fe-46f2-a834-5e82891d9f13 · outbound

This paper cites Building trust in clinical llms: Bias analysis and dataset transparency.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Building trust in clinical llms: Bias analysis and dataset transparency

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.289794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:23628d5050de0a783e95852d8a09af637c7fceb35e763242ee037f19fc69cce8

Observation f58203d7-0612-46ac-beab-f7a969ec5519 · outbound

This paper cites Eq-bench creative writing benchmark v3.https://github.com/EQ-bench/ creative-writing-bench.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Eq-bench creative writing benchmark v3.https://github.com/EQ-bench/ creative-writing-bench

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.368160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ab5fab125e63fcb735592ea70941162c70c325fe2e68750df6070b3b7de7d027

Observation 87aef807-ba0a-4b09-9a75-1a2b97409d78 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.294895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e97675d100c3a998e90eaebbf4920d6f31b6b8fa99c9ca082179e308837056a3

Observation ecc7d154-0fe6-434c-b2ec-717878a2d5df · outbound

This paper cites van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:21:15.039504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:241443df4ee9d34858b009f6414c14257d8efd1cfc15e6e2875b035ac480a568

Observation a9194573-f9b1-4726-a93e-d20276e60934 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:21:15.477936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:828d8c7f0db62329cbdc47f25005a8ef29658576628bb693160b67d2b69fc5c1

Observation b906ade4-0446-4b02-979d-2b8156a451d0 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.299211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:701c216eacbbd3e787a6b0011db3b84d1002212dcea5211bb26082896775c381

Observation 3cf4977f-485d-41f4-aaa5-3b550ffc7f41 · outbound

This paper cites Sutherland.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Sutherland

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.351536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:bb339192d4dade4a53d6f0d0b92ada24bc09d99b88127b3ddd8a071855f7ffd8

Observation b446829e-512a-4c08-b734-14fa0f269c19 · outbound

This paper cites Karl: Knowledge agentsvial reinforcement learning.arXiv preprint.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Karl: Knowledge agentsvial reinforcement learning.arXiv preprint

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.049063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:9459b2ff1c3492947dd5d05495d1464bc81bb39da5b429c8fafa785ca7dd5b00

Observation baea06cd-7924-41b9-9015-7b0e5ff2c439 · outbound

This paper cites Online rubrics elicitation from pairwise comparisons.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Online rubrics elicitation from pairwise comparisons

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.015919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:32802e5fa8c5dd4831136b6a1ccfcb17741037c8915cbd757e0846d5130932b4

Observation fcba67ff-a40e-41c0-a2c7-7c2a4c532907 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Proximal Policy Optimization Algorithms

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.128690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fed0d02774b01328a90c448ae8d19536e346da42eb3b039a9c4b9feaeed719b0

Observation 8f0b2e61-a066-4b38-b789-48f668c93a99 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:25da56724771c83e7346ae3ecc2860cf48fe69fc89e072958368a67534a4b4b6

Observation 0fdb52c6-4825-4dc1-bab5-1f2408888556 · outbound

This paper cites v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.655944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:6e78e0da4cb2b31fbf5fc0a6204e20b0aa7d23354b5bef365904bd313b5f8083

Observation 882aaab8-7a17-4de6-8f88-86d8b69a5d45 · outbound

This paper cites Book titles and abstracts.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Book titles and abstracts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.345305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:92758de1c76fba37e2542edcf225b69fea99b783607d99b06512973fb5d0671d

Observation 004af272-0f62-4965-88bb-d0c37e4a3a2b · outbound

This paper cites Mind the gap: Examining the self-improvement capabilities of large language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Mind the gap: Examining the self-improvement capabilities of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.278955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d678f5cde8b2ae61ee22e840e9a1b202b3842867d05da9487ced80ef9d1aeb33

Observation a16b9864-e86d-454b-844b-7afae12100a8 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Understanding the performance gap between online and offline alignment algorithms

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.758349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:45d266a76457f0e5156751954f4f94f53f37821ca694295a5789d17f695a9d0a

Observation d3c19b3e-ed8b-411a-8b86-303893dc19c8 · outbound

This paper cites GPT-4 Technical Report.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text GPT-4 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.585101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a0dfa0d374812725e4796bbdcd64e4392481fb83e3730254101f0dda922d145e

Observation 2d8c22e1-6318-4674-b612-7b551211262f · outbound

This paper cites Qwen2.5 Technical Report.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Qwen2.5 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.816751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fe94b85d35211ee4751b3b98edc4d71b5d9116368d9b7dec703692ba10d23659

Observation 181960ab-b1c2-485c-8966-4aba5b9046ab · outbound

This paper cites Will we run out of data? limits of llm scaling based on human-generated data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Will we run out of data? limits of llm scaling based on human-generated data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.340812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fbc0cf533661770390c5c8df97ccac244303e240da89447484dc96549efd4165

Observation eb4e54e2-abdf-414b-b079-7117811a630e · outbound

This paper cites Checklists are better than reward models for aligning language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Checklists are better than reward models for aligning language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.338244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:37dc6e5f8182fdd26ec695d56de08ded32a2a77a74ce9e36f395a80dc63be3ef

Observation f2bd45f3-8fbb-42a1-912d-b190b80ed32e · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.276123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c47c0518455ea1bb00abb465f7faa3a5f2c6c1b328d35c98ae318c6249d3ac59

Observation c0593d53-32db-40ea-8e65-eccbeecc140f · outbound

This paper cites Self-rewarding language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-rewarding language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.281802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b62c7b8b6e20da0df0343c974262a9980114fb3148ea70ba5192fdf1a78dfb33

Observation fdfe0c9d-e0e2-414b-9b12-156e970004f7 · outbound

This paper cites Better llm reasoning via dual-play.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Better llm reasoning via dual-play

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.500054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c4073c88d96eaa01adbcdc9d9280208fef62d3582ce36fac866aac6dfa36074d

Observation 26d16d03-9b4c-430a-b486-f55d7916c973 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:23:09.397455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:99b052771a2fca73565d7302b1392beb3318f6fd009a49d3784ca11a79e88b20

Observation d22ca3be-823c-4c42-84d8-f51a07a433b9 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Xing, Hao Zhang, Joseph E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.284299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:1f92d3dd80768b1831479f50d5158f994a930d4c425f4251ed03ffc1cc8fded9

Observation c794d79d-057f-49d1-ad6a-e9b15a81eba8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.695019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:6f12a820af19045d198e7d0fd6eabf688a829a5952e2540090f1a1c792b726a1

Observation c8b42532-f4e1-44ed-84cc-8cc5bb2445ca · outbound

This paper cites Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:23:17.018487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:49f89d8507fe525fa9a9b8c203553b722d14bcb9fc2dd40620fbc867f6a06b2b

Observation ec11e069-6aa6-4de0-9353-d988fd8028ef · outbound

This paper cites Self-Challenging Language Model Agents.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-Challenging Language Model Agents

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T13:21:15.160379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:93bac2e991dd1c9c93631c3df739d16f35fa184ec41707d9d2dbf44023484312

Observation 8e7bbc34-6935-49cc-83fd-12c915292c90 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.271410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:9fd1da3dc272f6d2db6626324ab730faf15e89bda343942db4a77bc8e8177ded

Observation 400b7898-0135-480e-9c81-e19e4b02b72c · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.359068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:6885fed77113587fcb4172341134311c8e96f6726a77b303e6692180587b0755

Observation b089ffea-c2bf-482c-ab3c-5383421a50ec · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.273495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:efa1767cd942293032ddff66ed15358f4b2040c02175b3f5d82413d50c4825a6

Observation 5383c140-1ba4-46a1-a067-85939b894728 · outbound

This paper cites Enzyme Identification.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Enzyme Identification

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.287023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:7c5fcd9a04eec18f9a45e2e56e4812cf43da436b33b7936c656aec81630c6be1

Observation 924c6ae8-f303-4138-932e-be8ec1d40f06 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.304236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d7975b5c80ecf72df0ead9b5ed88f47054a59cf1416dafd95c8c18837be9f3c1

Observation f72c0fbb-70ca-4011-b8f8-c017c65e93d4 · outbound

This paper cites Figure 44: Query (Instruction Following).

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Figure 44: Query (Instruction Following)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.325189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:986ef7d7826fbe087b9aa8aaf85bb8d14a4954727b7e02c9741ac2c13af3615f

Observation 1feccc5b-73fa-4e8e-950f-4968331f9809 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.292393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:f1d9bcdc29e6cc212e3f4abfee76060b73187cc53b5b19ac3384d1315ce9b6f2

Observation c7a45f16-28bc-4b44-b3bb-50d36e1f2d88 · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.301430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b1799494fd705d1daf49b2c74c5053de816f0a4aa378b792687a094b9c9d1ae4

Observation 2a659855-aabb-49c3-bad4-a3e3de8d1645 · outbound

This paper cites He stated that he would continue to work within the framework of the Philippine constitution and existing governmental structures.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text He stated that he would continue to work within the framework of the Philippine constitution and existing governmental structures

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.311963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:563e9a20c74f22f386e286763d934a39b7be96b1d18854b22d1f7e6b0ec88c82

Observation bea8f7a6-a5f0-47e3-a56c-49d776582e49 · outbound

This paper cites name": "Correctness of the first part of the answer.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text name": "Correctness of the first part of the answer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.266993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:1e224293ef70b21d31b28d29fad24cd8bddd7eea2dee524e6455e9c30282f593

Observation 4c7e9617-7256-4cd4-88ef-a583aebe27bc · outbound

This paper cites He stated that the declaration was made under duress due to threats from opposing groups and to protect national interest.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text He stated that the declaration was made under duress due to threats from opposing groups and to protect national interest

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.348055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:9f209d8e163c3b66d4d6855b05d8ae3c7d634e0459e8c733c5ac30d4cd294457

Observation abfbb882-59a3-4b30-b198-13856b5b3055 · outbound

This paper cites name": "Correctness of the first part of the answer.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text name": "Correctness of the first part of the answer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.384418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a9b76cc26f5e2054a7155be5fa57244c7f8f7ffc20274f319ea1522d4ec54e81

Observation e573b348-c5b5-4967-8abb-f7271f4bf106 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.262169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fda43d28cbdb3e66f0fcf67e65944effb54e8b4b4f24ec2e1ff43aacd7ee89b3

Observation c6c0fe7b-fe19-416e-a416-8d538ddf34da · outbound

This paper cites Ensure that it is necessary and sufficient to use these facts to derive the correct answer to the problem.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Ensure that it is necessary and sufficient to use these facts to derive the correct answer to the problem

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.382216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:4417a8dde6fab4a4c493eec098994b08d6a5be68396fa285955255ebb7f623c3

Observation b908ea69-d599-4d05-a49e-8de5f4d22450 · outbound

This paper cites * Provide a reference answer within <answer>...</answer> tags.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text * Provide a reference answer within <answer>...</answer> tags

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.255746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:4cae63881ede6a87e4bc88a033b72c2f284f5744318631c4e71502216fdc0032

Observation a32a39af-36e9-4b1e-8309-4265a1689059 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.259907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b7b792831911c44682490a24383a16db77cc630dd70b5fc3a04b7b912dd77bc9

Observation ef5d8dc3-4013-404c-941c-8e7f960e1109 · outbound

This paper cites ## Format * Enclose the question statement within <problem>...</problem> tags.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text ## Format * Enclose the question statement within <problem>...</problem> tags

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.264469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ca1bba47a44d9fdef6760d6afae7d89d64bc93a6a69e381118f8ce847493f501

Observation 40fc4e96-66de-4203-b52b-5e63518f6c62 · outbound

This paper cites according to the text sample.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text according to the text sample

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.330038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e2d308559041fc3e6a8a342fdc29224edf248f00b9bcee7fc92855438d477593

Observation 8af7f8f5-6931-4834-a013-b780d491fdc4 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.363704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:cf085c7bd3a1ee7465fe4464a28027001c35656984bcb1edcb9a2bf147a68cda

Observation 9ea86125-3e8a-4cca-893d-2c0ad75b6d66 · outbound

This paper cites Length: 1000 words.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Length: 1000 words

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.361507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:4859ba3bdce2b7810f9e35205f8c202ced24d8055969b884f18e02076e413bdb

Observation 4134d31f-b407-4820-a9b3-136d3806fb2b · outbound

This paper cites according to the knowledge.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text according to the knowledge

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.365723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:2faa4f6a432979fcc60dd6bf4d9e16a29947f358f53474550bdf54d4e8fc337e

Observation d778a6bd-ac8f-48d8-9a0b-f68d07aa19fb · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.386631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d68a15b71a4ccf89af03659b72f7951974e133619bf4b4042d8c93ff02b0fab4

Observation bf01da50-c8c1-499b-8678-5d123539d64e · outbound

This paper cites In general, the reference answer should have a high quality compared to the candidate answers, but this is not always true.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text In general, the reference answer should have a high quality compared to the candidate answers, but this is not always true

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.379607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b891c0341468cfa8cd5927c01688d651bf772b30fe44c04245451fc5761f7e00

Observation 994a32f2-05a2-41e0-8a3f-2c3120e3b373 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.372941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:63be624d6a8033404c11eb83d66191b61e0317cc2f0054a76884942d964843db

Observation a7f4a737-b37d-4a43-80c2-08d93a67021d · outbound

This paper cites factuality.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text factuality

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.269116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:2d52f0e0bcd4c121a029a0b6331e1e41f86ac63995fd32e8fb3b3ca4116529c5

Observation e54c8d39-e205-41e6-98bf-944ee822c9c0 · outbound

This paper cites YYY" of.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text YYY" of

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.327410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:80d15003dec4e04add8980fa42e525d934646d04f590ef4ff5835edc25e50f0d

Observation 97e0a20b-77e7-4aca-ad10-aece5e24e217 · outbound

This paper cites Not applicable.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Not applicable

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.370694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:8cab921a3fe44dacbbb4a183a60d80aeb11086a11229356e91bed1d4e2595899

Observation ccd96c6c-c2a5-401b-8cbe-6ffd1ce1abf2 · outbound

This paper cites XXX". To describe it, instead of saying.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text XXX". To describe it, instead of saying

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.375045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:82125f01230828eb5466b4b9378685ece778965addc93fee379f581d1d51fe2f

Observation e80b7bcd-7cdf-42d8-bdbd-e244093b742c · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.377227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:451fe7a8c43edf80c7bfb0a6d21be1755597da6e85cc5b01e862e88d2a4d10c6

Pith citing papers

Observation cbeba9fc-2602-4817-898f-741f2802f596 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.486944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.486944Z digest=sha256:543689388302406e35ab94927fcaeef957b888e50be79bb8768b9f14d6cc021a

Observation b4fc8f56-bfcf-4177-9ee4-d60d00d5ce5a · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:18.868835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:18.868835Z digest=sha256:5e6e8fc0738a1cd5e9b2c698a59777a82993beb3b7afd08bb76755d47f8acba7