Pith. sign in

Paper Citation Record · LEDGER

RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2312.16132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16132 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:03.233374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:56.284609Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2df884a-9610-4c1b-bfab-dc40fc85d249 · inbound

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning cites this paper.

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:03.233374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:03.233374Z digest=sha256:b0a021c2a2c2cd05d5a3902a3dd0a9d215fc298da8c7007d15d1e8604b8806a7

Observation e53bf8d8-ee54-47f2-8e0a-8de62ac2f943 · inbound

Personalized LLM for Generating Customized Responses to the Same Query from Different Users cites this paper.

Personalized LLM for Generating Customized Responses to the Same Query from Different Users RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:44:37.111621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:44:37.111621Z digest=sha256:386174927e0c17a30c61f48124a21e88d36f1859e46dd192e79cf7261147c4e2

Observation 97e458c5-a92b-4eaf-b794-1b648663accb · inbound

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs cites this paper.

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:22:38.369848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T06:20:49.096900Z digest=sha256:8d977fa7e8a7b3438978f945bd82b85d6900db925ef385633a2f02eae6f55e97

Observation 8b23cddf-ee0a-4dcb-bb08-e87e84c8344b · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:746fb9f85aabe4b31273f871a454bc7d38483a88278871c2bf06ca52a79152b2

Observation a0c3796f-39e3-4af2-bec2-537b97f9b1ed · inbound

BOOKMARKS: Efficient Active Storyline Memory for Role-playing cites this paper.

BOOKMARKS: Efficient Active Storyline Memory for Role-playing RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:55:03.885367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T04:51:44.394368Z digest=sha256:8c457212137562ea0b8db94cccbeba6d33da84df5889accd3ce75d8de587676d

Observation 6ff4c808-cd80-4f05-b07a-059a958857c8 · inbound

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models cites this paper.

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:48:32.102867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T15:45:21.524760Z digest=sha256:ed5636de8395ecac346f99c7efddb6645948de9714880720773f9efc0d69d4f9

Observation 70f0897c-49d0-4842-910e-b8a036789cce · inbound

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents cites this paper.

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.002713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T15:00:57.851111Z digest=sha256:9820d6c52f1eb1ee3773a45b42b2c342409a075af30d89620372242906d77814

Observation fb08e66e-b89c-4154-a6c3-36c9af69f4cb · inbound

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? cites this paper.

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:11:04.565958Z digest=sha256:19b6c76599dee2db1fb606feb875a5afda42dd25b4b954d7b7cfa63ab7692c19

Observation 9efd854c-5cd3-4898-94e5-3f505bd59384 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.704756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.704756Z digest=sha256:99abd3b6dc6e13707370c3ca567c91c7663486a2ec08bdfe7d3e2648a5c293a2