Pith. sign in

REVIEW 2 cited by

RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11387 v1 pith:N2CGLFXQ submitted 2025-02-17 cs.CL

classification cs.CL
keywords role-playingrolemrcinstruction-followingrolecapabilitiesfine-grainedinstructionsllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Role-playing is important for Large Language Models (LLMs) to follow diverse instructions while maintaining role identity and the role's pre-defined ability limits. Existing role-playing datasets mostly contribute to controlling role style and knowledge boundaries, but overlook role-playing in instruction-following scenarios. We introduce a fine-grained role-playing and instruction-following composite benchmark, named RoleMRC, including: (1) Multi-turn dialogues between ideal roles and humans, including free chats or discussions upon given passages; (2) Role-playing machine reading comprehension, involving response, refusal, and attempts according to passage answerability and role ability; (3) More complex scenarios with nested, multi-turn and prioritized instructions. The final RoleMRC features a 10.2k role profile meta-pool, 37.9k well-synthesized role-playing instructions, and 1.4k testing samples. We develop a pipeline to quantitatively evaluate the fine-grained role-playing and instruction-following capabilities of several mainstream LLMs, as well as models that are fine-tuned on our data. Moreover, cross-evaluation on external role-playing datasets confirms that models fine-tuned on RoleMRC enhances instruction-following without compromising general role-playing and reasoning capabilities. We also probe the neural-level activation maps of different capabilities over post-tuned LLMs. Access to our RoleMRC, RoleMRC-mix and Codes: https://github.com/LuJunru/RoleMRC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Activation Editing for Reliable Instruction Following in Narratives

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An unsupervised SAE-based method that localizes and adjusts instruction-relevant neurons improves instruction adherence and reduces refusals on a new 1,212-example narrative benchmark.

  2. VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

    cs.CL 2025-09 reject novelty 5.0 of 10

    A 65.6-hour movie-dialogue benchmark for spoken role-playing agents, with an evaluation framework whose main judge is also an evaluated model.

Pith tools