Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:18:10.028370Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2505.10218.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:18:10.028370Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:46:28.467404Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T04:55:03.945721Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6cfae966-dffb-4d05-8e94-ce1c626649e6 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241e77d0-f74f-474a-85db-982f1e3b23c7 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Roleinteract: Evaluating the social interaction of role-playing agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 882cadf2-4a5a-440d-9003-d517a3a42fad · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Socialbench: Sociality evaluation of role-playing conversational agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7c38c5c-6bd4-4dce-a658-c031fc41c446 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf980b2-05b0-4e5c-86b5-76abbe33944c · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Reasoning Does Not Necessarily Improve Role-Playing Ability
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5113a5c5-8b59-4378-af95-ccda4621e9de · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb372f5-2999-4258-becd-779dd5c0c40a · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward GPT-4o System Card
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f654fb-9a5c-4a86-960f-42d553e9899a · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Large language models are superpo- sitions of all characters: Attaining arbitrary role-play via self-alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 138d443d-b344-4929-b807-ba5f4802b8eb · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13398939-d52a-40c4-8bbf-2afd42b6a4bc · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Openai o1 system card, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8ff9a4-10a8-496e-a5f4-1e3923435d83 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c144e4-3201-457a-8be5-892a956b4d95 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CharacterEval: A Chinese benchmark for role-playing conversational agent evaluation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82816a0a-8414-4e29-8cdb-2fcf3bdbdad1 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Iteratively Prompt Pre-trained Language Models for Chain of Thought
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754e6b48-2c29-4a3a-b620-a3336e190f4b · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Rolellm: Benchmarking, eliciting, and en- hancing role-playing abilities of large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a780a87-d74c-493c-ab36-c88c54f04c88 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Chain-of-thought prompting elicits reasoning in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42dfb13a-28bc-4830-a64b-4adc55052b78 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward RAIDEN benchmark: Evaluating role-playing conversational agents with measurement-driven custom dialogues
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a280b2c5-cd82-4f8a-9425-13690aa9e951 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Baichuan 2: Open Large-scale Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1a2d6d-b117-4cf3-bb39-2e00c58fbdbd · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Qwen2.5 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2c8a2b-f5d5-47d0-bbe2-2c45c6e666e2 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c23b5b-aa18-47dc-a49e-38136c9fb685 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0871ec6-0677-4c77-87bf-3232a6aced83 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab06881-23f0-401c-b1ee-c1c623356297 · outbound
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward Characterbench: Benchmarking character customization of large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a8504a78-696c-40b7-9f00-5fd3fcf73866 · inbound
LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4bcfaa-181e-43da-aa89-d47fc9b5bf6f · inbound
BOOKMARKS: Efficient Active Storyline Memory for Role-playing RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.