Pith. sign in

Paper Citation Record · LEDGER

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.15068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15068 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:48:47.919007Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59633c9a-84b7-4b34-bf4d-4b5148322299 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.792003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.792003Z digest=sha256:f3d5339f992cdbd20bde9ef55d72123a51831e2a5394b66198a57fc42cb751b5

Observation 42880985-535d-434f-8c33-503a6db4456e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.824662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.824662Z digest=sha256:201c8ef20130e55de4a2da4b7299afcfc6f42617a4d25ee95f76c2ec09d484bc

Observation 82c7b8bb-6cd8-4b10-9960-85dbb0006565 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.829718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.829718Z digest=sha256:3c6ba9c6893379293895f61ba8a31a622f78d71c6bf397f7578c95cd2ccdf3d6

Observation 9b0f817d-33ce-49ea-a5df-f73e5b90ac41 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.833859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.833859Z digest=sha256:e7506b44a8b9f120602dd42b484f5c65edb821057bb5bc386c0067f52942f821

Observation 7cfdd1aa-03dc-4f08-9ddd-6574288410bf · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.838048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.838048Z digest=sha256:adedd7e77432c9255b7a31590ec8773f60a1399956b527403f677c8948669184

Observation 7b1ddaaf-0d5a-433e-a3ee-d08404b52603 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-15T19:48:47.956081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:46.842925Z digest=sha256:d3e48c667031fb87ad036495701ebd475e28d4f8697eac53a9c6bc336b4ccfca

Observation c5652fe0-37c2-4443-9162-5b0b1bb0ad3d · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.903746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.903746Z digest=sha256:08d18340df8764247d7c772f4168c5e06ad54051e3017ee742af26dfa4e513e2

Observation ab4f6210-4423-471c-a539-eaf97fb7ffdd · outbound

This paper cites A Closer Look into Automatic Evaluation Using Large Language Models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.963076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.963076Z digest=sha256:e7e4800ecdb4a06d5d3fabdea85589228e7ee4df1417bd69aefa1af8be2cd8a1

Observation b38cd49c-ff9f-41cb-9bb8-dd8aa9d367b9 · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.967870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.967870Z digest=sha256:058d3b78f9d8cfe708676dae6fbb8ca78ee54cf9e60bec87b3c937c7957d564a

Observation 70d93744-92bd-46a9-a678-06b9312ab7dd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.972385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.972385Z digest=sha256:df5dcd4fb64aed421d481d6a4f154da9757db640e58ac0bd39823cb9e7bf7df1

Observation 9fa16c0b-d5f1-4873-8484-7c74e090c983 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.976457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.976457Z digest=sha256:1967f63799875eaf2778ab93d19c72fbecf0d916539097f68b901b6571333573

Observation 7a0908a3-b4e9-4944-8c41-2c9f90b74268 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.981211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.981211Z digest=sha256:dc4beabf5c9b49cef85c9959832ad5f4f58d38945ec98fa8509ecff9d43e96e9

Observation c711cc16-073e-493f-a9dc-5fd00a35fd2f · outbound

This paper cites SummEval: Re-evaluating Summarization Evaluation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation SummEval: Re-evaluating Summarization Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.985186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.985186Z digest=sha256:2f80cef07fa71f681dcdcd4c051c7040ee4ca1ff8ec0353606d0a98edef952ee

Observation f9f89c67-2d0e-487b-8ab7-d9944da35723 · outbound

This paper cites ELI5: Long Form Question Answering.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation ELI5: Long Form Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.993204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.993204Z digest=sha256:1e68283bfb5a9281161d464752a1190539ea306d0f2e84c2458d879c3022852d

Observation f87806cf-b748-4071-8717-7ddfc286289e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey on LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.109728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.109728Z digest=sha256:fcb6a1c478ea8c958f40836affeccd7c2fb6b243b960ba80c9dc51a8999705d8

Observation 4a4a841e-ca3a-4ddc-aa54-1cafb9661be7 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.152322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.152322Z digest=sha256:fb903238d967241a0d6ef76056f1db6467b373613e9cdeb63aba87d1e2c462cb

Observation efa151a7-b8e1-4f0f-80bf-09e365e4def7 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.156442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.156442Z digest=sha256:6073cfe2ee510e67829d663f9c80f4ae3de971262cb7f56f5df8ab938b0e0c44

Observation 4c742731-11ea-453a-b364-66c2639d2694 · outbound

This paper cites Hurdles to Progress in Long-form Question Answering.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hurdles to Progress in Long-form Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.160503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.160503Z digest=sha256:52189ba1a4afcdd01008ff23d4722f5b1631c4ba5f266b52499ab5a5beaba022

Observation 07198add-f6ef-4dbe-a58c-f159f398f067 · outbound

This paper cites No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.164830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.164830Z digest=sha256:099d710b2cd1b5e15140b2fe9539e144beb41de08ee71f7f82f61e23266e0d9f

Observation d7fa89f0-82dc-4452-b104-3b6221a3f719 · outbound

This paper cites Kullback and R.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Kullback and R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.169094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.169094Z digest=sha256:ccca060bdc3b4bb278f5a2a1af5a8855a2f5b373f11a7b226b16a7de74405c49

Observation b19010c5-7788-4fde-bea5-fb3b124d0b0f · outbound

This paper cites LongForm: Effective Instruction Tuning with Reverse Instructions.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LongForm: Effective Instruction Tuning with Reverse Instructions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.173159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.173159Z digest=sha256:b3c9f3bdd447a5046535258b6eeac408ee129f0507766b3ffb55ec9281fbc758

Observation b1ada493-ff02-4dd9-a496-6bf5f803497c · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation RewardBench: Evaluating Reward Models for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.177700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.177700Z digest=sha256:b98eea12986af292dddd5d8cf6aebb8eed060bd1ab9c744582f39b097e7c7253

Observation 911a17c8-de7d-42c1-a6ef-a86504782f47 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.196869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.196869Z digest=sha256:e4e0ac41a11fd1e05187e2b4fda1dd21ac37f4f241c6323e234c0487dd61da16

Observation bb7991a5-270e-4e7d-b946-c45ccf307c73 · outbound

This paper cites Deep Reinforcement Learning for Dialogue Generation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Deep Reinforcement Learning for Dialogue Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.266034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.266034Z digest=sha256:a760cbe9e1aa7f985a3312de9608e44c1f3d85be59550d98891b706ed83b898d

Observation ffc1a410-d56b-412f-9094-dce7e371e508 · outbound

This paper cites PEDANTS: Cheap but Effective and Interpretable Answer Equivalence.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation PEDANTS: Cheap but Effective and Interpretable Answer Equivalence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.341857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.341857Z digest=sha256:f76eeba9f85602f25dd3f64d134832f6e8bc7b9b4360bdd5dccc716a2b3abdaf

Observation 49f107d5-ae54-4120-9047-51de611f1b45 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.072381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.346440Z digest=sha256:c4a36cb86b71180775c72bb4fff23fc6fa00b09f9d7ad3922ea2c1b6225a2cec

Observation df51d059-bcb6-4a3b-bbcb-72dac589f17c · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.349856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.349856Z digest=sha256:5ba2af55e9a516212adc2d84b21cbaffc40f523a4b6873b4ae27834db5f43a64

Observation 8c56b692-c4c2-4d25-8589-c10ec768514b · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.060450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.354409Z digest=sha256:d98ae7063d6e5730864d4c01a537a6cb4f09d8054864b750f74e9e2a37e3da54

Observation 6e17e919-6fd1-41e5-8794-8c9214c39b57 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.358973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.358973Z digest=sha256:686a94225c34283be27ad15bf1112acbfde143fd92df629a82ad7ff915e61aa6

Observation 4291197c-ead8-4f98-9239-b0b637b29654 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.363358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.363358Z digest=sha256:188e37bc528daeff63a15a0e2a5483eb9aabaafab30563a22da9f28a652a6843

Observation 8833ad02-2bce-4a3f-86a6-d04de55c472b · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.368219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.368219Z digest=sha256:cb446dd38725d4cc5937d98c3eefe2f339b0cee6689ecd3e3d1e916aa53c0ff1

Observation 7b1cb00b-6330-4a72-8ca3-e3c80af6e88a · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.426148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.426148Z digest=sha256:d0cd8a5ee7470ead9494467ff7a8d9a9d27a98e7586ce1219ad91ef3dc3cee73

Observation caf6aad0-cb9e-486c-90a8-55ce572490db · outbound

This paper cites Qwen2.5 Technical Report.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Qwen2.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.445193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.445193Z digest=sha256:c562150ea7a6e920cb032144fc0e4fe4307fd67de9929ba4fffde4976742f8ae

Observation e0d3f46b-16be-4087-a909-0c6928e5e594 · outbound

This paper cites Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.490328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.490328Z digest=sha256:9f3ff5d746e8cec565bb35d0e030cec310b4b5bc68fefee27663208e39474619

Observation 2172f24f-836d-4973-bc8c-088284703ebc · outbound

This paper cites Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.494498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.494498Z digest=sha256:21bb913552511bfb946f45e7fff585aec7506231c4983876fda471713d898b40

Observation 0ed25aac-2353-47cb-8753-c3843f9b9a99 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.498566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.498566Z digest=sha256:c10e43d9d7c5f9289aed8dcbd7b2bc55ab34cd4fcd90f34ecb9ea0156bf3b41f

Observation d8c6387f-e736-4847-8856-3b8f8f9b70ef · outbound

This paper cites A Survey of Deep Reinforcement Learning in Video Games.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey of Deep Reinforcement Learning in Video Games

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.502875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.502875Z digest=sha256:5d158fcbbce27392185a4740c507159a48b3ffd81321fa003bc23b70350151e6

Observation 21355d6d-53d0-42d9-89ce-3933300a6fc7 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.043155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.508657Z digest=sha256:2e8c449db6bc29c55e4d2effc1a3f39aa8d7025d968a24232ba940e28e0a7a81

Observation 3343d9bf-eafe-4e63-80b1-18cb54921874 · outbound

This paper cites Learning to summarize from human feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Learning to summarize from human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.511830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.511830Z digest=sha256:e0dd4cb1eac4ff6a8bc26be3bcf3d79294d772077b9f53cb9f4a2a72cd9bd9c6

Observation 398115cb-a38d-4ff8-9663-c0d3013a26f7 · outbound

This paper cites Hashimoto.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:48:49.030890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.515624Z digest=sha256:d1f39adcff81125c2fbd51865d81e19f373500a5b503e38dc23861bd4be3eee6

Observation 72d469f1-4e11-40b5-8359-e58a2bcfc064 · outbound

This paper cites Hashimoto.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.518693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.518693Z digest=sha256:80f6af3449c5ccc030d3759b1feeaf19cafe0f4e75c0e98df601a92c4be622ef

Observation 412274bc-b360-4114-9db9-a1b5c0282261 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.522256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.522256Z digest=sha256:0ca6c5f5c9581705a54669fc350324ddfd435bdf0f4c15f593eece3d5952b68f

Observation 11899969-0add-487d-bb90-380a48fbdb83 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.589286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.589286Z digest=sha256:d45b39c29ec017eca69c116919a1a0c7a255cee5611b41ec485512551858425d

Observation 92417eca-0019-483a-90f9-b093826c8ad8 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:48.945108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.646879Z digest=sha256:f2a9d0fbdba7754025aadc528b55f0357a153cbc030815f154b830a391aa9aa5

Observation 290d8a8f-01fb-4548-89df-a767ace4b939 · outbound

This paper cites Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.650992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.650992Z digest=sha256:59110bcee03da0903b1ab27c293d68366667d6026c7a37f8c95ebf68840dd1ae

Observation a98bc16a-ab2f-4290-95e6-7b81c94ec08d · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.655647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.655647Z digest=sha256:a360fadbaf9eff3f186c55fddc885bd340c90e03bfe6a9c5589b4345d06cc627

Observation c51569b5-6d49-412f-baa6-70af2a9c25f0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.659845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.659845Z digest=sha256:f82052c32add437f77ae67a2aba4cdec35b2566efc05b93de41a6dae71a78e7f

Observation 5608c357-1103-4bb8-950f-b36f355de604 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.664012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.664012Z digest=sha256:344897dec190b7d504139d8a99dbd1c2ade38cf2a68699acbd52c3d54e7daacd

Observation 9e215979-2aa4-4017-9edc-3cc08fc24453 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation BERTScore: Evaluating Text Generation with BERT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.667853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.667853Z digest=sha256:9f1f1cb08e004339439ad66f5d5b75ab795892377d8114699aa6d0342a2bb649

Observation 9ea4d148-32eb-46c8-83ca-a04a4bf37f9e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.671746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.671746Z digest=sha256:8e1b5fd274dc82b5ab42a7ce5d8a21d95db44d811a99655b8731f99279fb4729

Observation 35925866-a29d-4310-a845-2157a35a91f0 · outbound

This paper cites LIMA: Less Is More for Alignment.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LIMA: Less Is More for Alignment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.676062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.676062Z digest=sha256:45364955d9b18ce23f6ff3422c93c6e063e7f1542d43a44893d72771068a555a

Observation 9d2223ba-ca32-4984-a71d-2faf3e8d0469 · outbound

This paper cites Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.680838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.680838Z digest=sha256:4031a696c994403cc7e580cb3fb7b9d1986efce669e7a50fcdcbc8fd2c19b412

Observation 51ad07ec-509b-42f6-98b4-1090207865f2 · outbound

This paper cites MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:48:48.164263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.684812Z digest=sha256:98ff9964b7fcb90a05e9f69a9410deb4d19c4197665218d7feb66a45986d6012

Observation 98ba6f26-631e-4f9b-8f52-c5861c9c830e · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.776043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.776043Z digest=sha256:5232b12f3ad3f289095083d21382adfb7493937c4dbe00d4482f42f8836ea44a

Observation a61909ed-061a-4b0e-a02b-c77979fd6e2f · outbound

This paper cites online" 'onlinestring :=.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.869226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.869226Z digest=sha256:701ab9bb5bf1c4750419b32aa49b2d849d296f7fff5265c04f5f254de545da3f

Observation 729d134a-ba3a-4338-a043-824819ecca29 · outbound

This paper cites write newline.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.919007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.919007Z digest=sha256:83b34977ff7ad65b8e8bfbf8bbae7b39cb145ffec352e9c66a8bdd3fecda902e

Pith citing papers

No inbound Pith citation observations are available.