Pith. sign in

Paper Citation Record · LEDGER

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2505.10218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10218 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:18:10.028370Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:46:28.467404Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T04:55:03.945721Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6cfae966-dffb-4d05-8e94-ce1c626649e6 · outbound

This paper cites GPT-4 Technical Report.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.938638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.938638Z digest=sha256:3e42327dacee41745126f40619be399339752ea96c50811cb7459b96680f0678

Observation 241e77d0-f74f-474a-85db-982f1e3b23c7 · outbound

This paper cites Roleinteract: Evaluating the social interaction of role-playing agents.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Roleinteract: Evaluating the social interaction of role-playing agents

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.300097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:09.943790Z digest=sha256:2a08d94baa89b3b00b483620fcc97a01ab136c73ac67d1d3cacb5638cceca80e

Observation 882cadf2-4a5a-440d-9003-d517a3a42fad · outbound

This paper cites Socialbench: Sociality evaluation of role-playing conversational agents.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Socialbench: Sociality evaluation of role-playing conversational agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.288389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:09.947642Z digest=sha256:3fb93a3a0b91792ea330119cd08feaabed084937dfcc5f741ad218a18031ef91

Observation c7c38c5c-6bd4-4dce-a658-c031fc41c446 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.952201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.952201Z digest=sha256:0336946c31ecc9982cdd3cbbecb35f8142eb039b750c5f7d5df92eed94d8bcbe

Observation 8bf980b2-05b0-4e5c-86b5-76abbe33944c · outbound

This paper cites Reasoning Does Not Necessarily Improve Role-Playing Ability.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Reasoning Does Not Necessarily Improve Role-Playing Ability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.956907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.956907Z digest=sha256:9c838b2a93b97b5c700c131757846ce43fb919c9c5657548394fb980e2ac73a5

Observation 5113a5c5-8b59-4378-af95-ccda4621e9de · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.962102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.962102Z digest=sha256:22ebd38d12d0f3ecd623428683e78a3166bf6bb40ac0122db3561f4e74c58ccf

Observation 9bb372f5-2999-4258-becd-779dd5c0c40a · outbound

This paper cites GPT-4o System Card.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.966533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.966533Z digest=sha256:b17e96daaf13303ec1c1c0065f2342b55bb6e84ea18fe31b3d113a56025f4500

Observation b1f654fb-9a5c-4a86-960f-42d553e9899a · outbound

This paper cites Large language models are superpo- sitions of all characters: Attaining arbitrary role-play via self-alignment.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Large language models are superpo- sitions of all characters: Attaining arbitrary role-play via self-alignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.269047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:09.971384Z digest=sha256:313c6d76e6ebc873259f60c1fd11fbe7f044a68981e749ae531654e4f15d9eef

Observation 138d443d-b344-4929-b807-ba5f4802b8eb · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.975019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.975019Z digest=sha256:4c3654a2c466011778cd58601545a8ec0cfd27ed38aa0c896af8a11ed190dc0a

Observation 13398939-d52a-40c4-8bbf-2afd42b6a4bc · outbound

This paper cites Openai o1 system card, 2024.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Openai o1 system card, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.978951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.978951Z digest=sha256:85a98c33cd21a6563682e621696efe044fe5d65db8e285509271e750866f2062

Observation 9c8ff9a4-10a8-496e-a5f4-1e3923435d83 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.982732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.982732Z digest=sha256:76421f9ba5550172a311a77d82bcedbfecebcca0a87626fa22a741e322af46ba

Observation 47c144e4-3201-457a-8be5-892a956b4d95 · outbound

This paper cites CharacterEval: A Chinese benchmark for role-playing conversational agent evaluation.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CharacterEval: A Chinese benchmark for role-playing conversational agent evaluation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.241822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:09.986964Z digest=sha256:8751a5a5e2145001e78bc245afd3ce714fcc8e21e11ab4c1a59fa92fcadac1b9

Observation 82816a0a-8414-4e29-8cdb-2fcf3bdbdad1 · outbound

This paper cites Iteratively Prompt Pre-trained Language Models for Chain of Thought.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Iteratively Prompt Pre-trained Language Models for Chain of Thought

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.991576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.991576Z digest=sha256:cf603da761c6fda4b6784a62e2847f50059469bee6c177fa833968e8d312fa17

Observation 754e6b48-2c29-4a3a-b620-a3336e190f4b · outbound

This paper cites Rolellm: Benchmarking, eliciting, and en- hancing role-playing abilities of large language models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Rolellm: Benchmarking, eliciting, and en- hancing role-playing abilities of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.229090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:09.996137Z digest=sha256:6ae496879da2d96412d43a04e069260942e78852d1b6a100a54ec5164bd7c925

Observation 4a780a87-d74c-493c-ab36-c88c54f04c88 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Chain-of-thought prompting elicits reasoning in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:09.999746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:09.999746Z digest=sha256:910acd1847c080f24f8e97c49812536b95dea67f57ba9ff02181ed7a424ec5fe

Observation 42dfb13a-28bc-4830-a64b-4adc55052b78 · outbound

This paper cites RAIDEN benchmark: Evaluating role-playing conversational agents with measurement-driven custom dialogues.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward RAIDEN benchmark: Evaluating role-playing conversational agents with measurement-driven custom dialogues

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.207370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:10.003816Z digest=sha256:a5657a0896f06a27f1cc1ba4b6bbb29d4442f5080a9b1f43fb38506b447021e1

Observation a280b2c5-cd82-4f8a-9425-13690aa9e951 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Baichuan 2: Open Large-scale Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:10.007927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:10.007927Z digest=sha256:0e5db909cf7f2c2a891b0ad92fcd26a09d521b3d386e96fc27ce1cb3a18f18bc

Observation 6b1a2d6d-b117-4cf3-bb39-2e00c58fbdbd · outbound

This paper cites Qwen2.5 Technical Report.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Qwen2.5 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:10.011855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:10.011855Z digest=sha256:727374e2afb9721d9d27d46b285dcddeeb013f4d32d04ce313f6f8d8301ed96e

Observation 3c2c8a2b-f5d5-47d0-bbe2-2c45c6e666e2 · outbound

This paper cites BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:10.015801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:10.015801Z digest=sha256:262b2c9959474a93f0555c99f4d1fa6edd441d55095749a4989ea8da6999ed37

Observation 12c23b5b-aa18-47dc-a49e-38136c9fb685 · outbound

This paper cites CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:10.020320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:10.020320Z digest=sha256:27437dc211be1c03b86b038ce80da89caa6f75bff2b0d61cf9a1fe103e0a2d4f

Observation a0871ec6-0677-4c77-87bf-3232a6aced83 · outbound

This paper cites CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:10.024091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:10.024091Z digest=sha256:de58a5476b0c74cec08e31d7d0a6fe47930a1720c1845d70a10bf9277212c088

Observation 2ab06881-23f0-401c-b1ee-c1c623356297 · outbound

This paper cites Characterbench: Benchmarking character customization of large language models.

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Characterbench: Benchmarking character customization of large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:18:10.194465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:18:10.028370Z digest=sha256:1f631366b5f83906184e1a3d2913cc30a69981ffd9fde6f24b1d239a03f1af1c

Pith citing papers

Observation a8504a78-696c-40b7-9f00-5fd3fcf73866 · inbound

LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing cites this paper.

LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:28.467404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:46:28.467404Z digest=sha256:de04337da2798625263e2a0af00a0e1d802a9b686adb2a0c8cd81408dd61f045

Observation 3a4bcfaa-181e-43da-aa89-d47fc9b5bf6f · inbound

BOOKMARKS: Efficient Active Storyline Memory for Role-playing cites this paper.

BOOKMARKS: Efficient Active Storyline Memory for Role-playing RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.949279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T04:51:44.394368Z digest=sha256:e146ccb5e702d21861bcd9bd69ab0fb7997bbe8b460fce795b2adc2753ae2a3d