Pith. sign in

Paper Citation Record · LEDGER

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2505.14106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14106 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:23.733585Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.701496Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:34.994786Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9861fdb-99de-45c2-8df0-515c31a77adb · outbound

This paper cites Conversational Health Agents: A Personalized LLM-Powered Agent Framework.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Conversational Health Agents: A Personalized LLM-Powered Agent Framework

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:18.519445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:18.519445Z digest=sha256:36166aefa46b25d8ac3ea11edbeeffc4f35ab47a0ba57e8d18b8cf6b06d6157e

Observation e00bf5c9-bc0c-47b8-b2f7-cd2eb10dc9d3 · outbound

This paper cites Persobench: Benchmarking personalized response generation in large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persobench: Benchmarking personalized response generation in large language models, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.054834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:18.625291Z digest=sha256:1d4fddbd1a01473b2cd7f5d5be31f1b60536a6c01c1ae977f7afa2ae38e700e6

Observation 03167942-a7e7-48c2-bf1e-e6ae533ea020 · outbound

This paper cites Claude 3.5 sonnet.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.909324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:18.729628Z digest=sha256:582b429e4d51e048e841213e43e36537f959de601746ae4b4ea60d94937859f5

Observation 237bf522-348b-4651-a528-b532151c0119 · outbound

This paper cites A Little Human Data Goes A Long Way.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Little Human Data Goes A Long Way

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:25.276669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:18.993776Z digest=sha256:f22015eac043faa077f6b056032235fb368264065eb10c4ecd5b80574072055f

Observation c3e01726-ba8e-4cf9-b027-26cfc2176d36 · outbound

This paper cites Personalized Graph-Based Retrieval for Large Language Models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Graph-Based Retrieval for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.113061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.113061Z digest=sha256:6203b5a57f773cc251225ceb7a037a4bd912f945de6a5b1256877e532214864c

Observation 03232c27-8236-4033-b7d8-f3d2f54d19af · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.244719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.244719Z digest=sha256:88250890f8c06804407e4c79feb5eacb45ae94fa73276dd49c3107f2b9044b64

Observation a903f797-26e0-48dc-b27f-1e35298c32d9 · outbound

This paper cites LoRe: Personalizing LLMs via Low-Rank Reward Modeling.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations LoRe: Personalizing LLMs via Low-Rank Reward Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.367275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.367275Z digest=sha256:1f7951b2deff76b1d155da4a76da79617ecf6c62a69d903c116c08c97f1b9063

Observation a8722fdf-21c2-468e-b36a-a02ebb7c345c · outbound

This paper cites Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.490084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:19.513223Z digest=sha256:27db8024d09284e5c0459b5d30f022a35f34e301ec3fe9a8142293cdf3fd8355

Observation 74b080f6-b2ed-42c0-a046-72636d8e93b4 · outbound

This paper cites Beyond prompts: Dy- namic conversational benchmarking of large language models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Beyond prompts: Dy- namic conversational benchmarking of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.268054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:19.655393Z digest=sha256:b18065d75945e1f0796f8f38d0aad144c6f1d8219275254424cdf9676196a675

Observation f3e6934d-15cb-4155-9c60-ad8fd8a2b49f · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.075713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:19.807578Z digest=sha256:7edcbb80993dcc677c078c225ccc5ce4fb7fdf48089b593d95b94cfaeee2a590

Observation 2ef064e0-e58f-4565-9410-fc52a2f1045c · outbound

This paper cites When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.894818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.894818Z digest=sha256:79f943bae9b11d3ff450f4557eedc49d44bce1add314bfd374e03063deb2cb18

Observation 3cc7d673-0b61-4547-aa25-30019abbd5a8 · outbound

This paper cites REALM: A Dataset of Real-World LLM Use Cases.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations REALM: A Dataset of Real-World LLM Use Cases

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.811302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.033370Z digest=sha256:7081adb6258478c51f73325cfbf6eff99f4a4b2edd89efa67c34f08c33352834

Observation 4e070871-557d-43a0-adea-d59997ee4ab3 · outbound

This paper cites The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.909265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.144923Z digest=sha256:3af5638fa47bbec530b45b957a9f2e647dcddc39da5ba0e472272c3ccce4ae58

Observation b06db013-27d0-4816-9cb6-68ce7003e4d0 · outbound

This paper cites The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.761772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.324986Z digest=sha256:33dde87e4460817d793103cf090b023ea12229d3ee97c716cce8b2ea3e24fc62

Observation 88126155-e729-4f8f-8fc5-f140559fa286 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.595175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.465431Z digest=sha256:64533f34317ef9ade2406db3998ed067fed7be527fb296d1bb822cdc2535d25f

Observation c48a18f2-6a9e-40d5-b11d-dd9a31c50a36 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations RedCaps: web-curated image-text data created by the people, for the people

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.580260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.580260Z digest=sha256:1cba1caa3c5f1dd2869051010d8973ed65b4a8382c5b27acc9af08626d17a4fd

Observation 9a06c4d6-7fac-4a12-a0aa-9fc9e604ac27 · outbound

This paper cites The llama 3 herd of models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The llama 3 herd of models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.662819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.662819Z digest=sha256:c592eff128d824ec784b2fb89ec05c386c452c3b7f32d07db02d77b38843f4d4

Observation 850ceb35-110c-4dc9-9b9a-a27a373e84d2 · outbound

This paper cites Ruddit: Norms of Offensiveness for English Reddit Comments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ruddit: Norms of Offensiveness for English Reddit Comments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.552148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.731343Z digest=sha256:4c03d785b2486e54120129f64003565e577d4bf565c2cdf0079c8049aa174e48

Observation 194d8e5f-4214-4e7e-bfc7-0a4e1ee518eb · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.382963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.867591Z digest=sha256:96f1dfbb9eee109e01e35f764e367cdf7dd9f21484a035b3ea8555c9c03c7e82

Observation e4f4b3bd-d184-45bd-9946-d970bf281d4c · outbound

This paper cites Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.165335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:20.984363Z digest=sha256:d5998e816bbdfdc3aa8d6981407d47f6c7bd6621e7f09d25bf0806f7a57ec172

Observation 570f0968-3ae8-4466-8fbe-45f11d61adc4 · outbound

This paper cites Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.936888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.117526Z digest=sha256:a5294edd7645b9162474a1026009944e8ae05983a8ba202b783c59a365b948b6

Observation b6daae78-9455-4ebb-ac01-1c48db55eb46 · outbound

This paper cites A framework for building adaptive intelligent virtual assistants.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A framework for building adaptive intelligent virtual assistants

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.672196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.243210Z digest=sha256:da5f2a5151048f79782558559571fc98d9f9e951951fb7a478b7c30872b2bb01

Observation 77ecc1ea-0684-4054-b46c-9b666cf43b3c · outbound

This paper cites Teach LLMs to Personalize -- An Approach inspired by Writing Education.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Teach LLMs to Personalize -- An Approach inspired by Writing Education

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.349336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.349336Z digest=sha256:72574d17863792bb64da111cbc781ce5f207fbcade611568a18cd27e16a20ebc

Observation 1945642b-fd5f-49bf-a9bd-77d49420dab2 · outbound

This paper cites Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.437751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.487031Z digest=sha256:75b26af28be06685330e1ee6a522dffbb7ced47c6f6a890ad8aba8b0928ca15b

Observation 62133eab-494d-409c-9d3b-59c8a7b733aa · outbound

This paper cites Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.240242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.584797Z digest=sha256:e5a863cdc9892f77dd03bb65df053df6c4e20e5bf827507ace73d5eb5194e42b

Observation 955fdb45-569b-4517-b7be-6a5169debecc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rouge: A package for automatic evaluation of summaries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.651056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.651056Z digest=sha256:4f6ee943c824be10d4bd8ebb2f6de7d118c246d6eebdd6b869fe2e9d63226f03

Observation 646fc09e-b860-47e2-a683-92459fac4bb5 · outbound

This paper cites Persona-sq: A personalized suggested question generation framework for real-world documents, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persona-sq: A personalized suggested question generation framework for real-world documents, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.971000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.760648Z digest=sha256:2490bd07805b73b7b90f3be3ee5db73ebd78f67ad11bce859d28f0358db19d1b

Observation 9533affa-216e-4632-8ac6-64befd694cc9 · outbound

This paper cites Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.829533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.906459Z digest=sha256:ce2e7f5401be82c7e9328dd676de4d9479097c6b423bfe695ed46a920364b2d1

Observation a054c5ee-571e-44a1-bd79-40f20b53da15 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Gpt-4o mini: advancing cost-efficient intelligence

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.643229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:21.975869Z digest=sha256:3d0de3d826aa36eff943d5006dae51ea9e380318f9ac3b412fb382866b0fd37a

Observation 6ac76a55-64bc-42c0-96ca-c0cfd0f0339c · outbound

This paper cites Introducing gpt-4.1 in the api.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Introducing gpt-4.1 in the api

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.480886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.043964Z digest=sha256:55b78fbf9aa540076db9db2b263b97f59001cee9f75bd600621d257ca0d87831

Observation 9d69a277-3a48-4798-855c-a90eab1feb4e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Bleu: a method for automatic evaluation of machine translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.172514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.172514Z digest=sha256:e6e3b79295e7b1b584a63bb349400a9517abbe8615fda4e822e10c64c872b75e

Observation 65e016f6-564b-4426-8085-d6a390ff8535 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:31.250804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.264169Z digest=sha256:8a608a580a2c930771971f7fab59d1981aaced50964b416798f9bcfbc29824fd

Observation 0fed4877-ccbd-4a02-9aee-37490e77aec9 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:30.988064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.394518Z digest=sha256:190350231dd796b2dc9f0466c0d69225d9ab71c5f451a344aaf8f411a0262bd7

Observation e44488dd-ecf2-4aaa-83fb-0f4bdac730a5 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.486570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.486570Z digest=sha256:ba42d01b39ee88444b638419fec6b27b394fb49fda98c37b393f8ad8e49ae5dd

Observation e08ac11f-6f2a-4ef2-bcd4-b722c05ab059 · outbound

This paper cites Lamp: When large language models meet personalization, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Lamp: When large language models meet personalization, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.780537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.581458Z digest=sha256:7c70d56325f65cb84661acd1a6aa545ff6be23391f122130ffa12b7c2397464a

Observation 639f9236-0fca-4285-9c56-1765c052d4e5 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.649086Z digest=sha256:c93dda2c42c1796edfc5d88406bf10d81958d7bddcda14ddd416a05a999fd817

Observation dbf2d392-76b7-4210-bacf-45ab18e9a5db · outbound

This paper cites Democra- tizing large language models via personalized parameter-efficient fine-tuning.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Democra- tizing large language models via personalized parameter-efficient fine-tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.623584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.762396Z digest=sha256:def1ae4929a89eedf44ddd9212b60bec16e9064b1e5f251f31b3bba2e92178cf

Observation 2d87413e-c6d6-43ab-9c83-59a6cddc2c4e · outbound

This paper cites An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.404821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:22.887964Z digest=sha256:60d1f9ee3fa6e361c2dd9cc94f7b45a1145487b93cdfebbd2aaff321363b232d

Observation 6b15c0d2-b97b-46a3-abd4-7495a23729a5 · outbound

This paper cites Position: Will we run out of data? limits of llm scaling based on human-generated data.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Position: Will we run out of data? limits of llm scaling based on human-generated data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.943347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.943347Z digest=sha256:3f834d02b94b888fcc159a77c828f0a0c7b5d7ed4bc545eac59435845ca5f5bd

Observation 2d03cefe-d484-4bf3-8557-0007abc5c064 · outbound

This paper cites Personalized Multimodal Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Multimodal Large Language Models: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.001354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.001354Z digest=sha256:0c54a7de7826d3c0ee0b428440b69de759f3e8f4c953beadb9a07d0578c0fad5

Observation 0c64fbc1-cfc5-4224-ba00-99c35c647e1c · outbound

This paper cites A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.118282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.118282Z digest=sha256:b016fbd1c25b66e4d4a53e07a9c41ee39cd5c222cd0ecd26e559fd4aa5a53831

Observation f5e06eb9-0d06-48ae-bf85-deb4ba1386b5 · outbound

This paper cites xdial-eval: A multilingual open-domain dialogue evaluation benchmark.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations xdial-eval: A multilingual open-domain dialogue evaluation benchmark

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.185319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:23.196806Z digest=sha256:c66d7fadab722c9e3b5f7975841e0bfe9215643a3fce10e52eaba821543a0e7c

Observation ca8f1e77-8842-440d-a565-21323465c189 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:25.986776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:23.314664Z digest=sha256:69b1f13d5f76f1d637c87710a96196611e7a02d4581ad858296c112f4ce87988

Observation e23151f7-29ff-40fe-9dbb-d849f8ae67e2 · outbound

This paper cites Personalization of Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalization of Large Language Models: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.443293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.443293Z digest=sha256:b5ca4cbbf315fd008e55c93393944082f9ae47ffa9941a9147ec7004df12ac09

Observation a9e526ef-fff8-4464-832b-32f5844900cb · outbound

This paper cites DiQAD: A benchmark dataset for open-domain dialogue quality assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations DiQAD: A benchmark dataset for open-domain dialogue quality assessment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.739429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:23.527996Z digest=sha256:df47ff0c85678a297e1f3905434395e8c21b713b7eec11bd092b65133ed3615e

Observation 791e650c-e121-426e-87d2-3373f14446e5 · outbound

This paper cites Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.537365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:23.662353Z digest=sha256:26176fca5671fd3fadd3b20a94f9b7c2d55d51b191124f66eef3b76f7c7cb906

Observation 7c921583-99e9-4d70-8560-649b95daa7cc · outbound

This paper cites Best Response.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Best Response

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:43:24.106045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:23.733585Z digest=sha256:21fc86631fe2be868502c90758beb95817560c525e03e7e417f1ab33cf53fc02

Observation 3c5c23a3-3e7f-4e88-9235-16ffc7f9185b · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:43:34.736747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:43:18.862204Z digest=sha256:f61439045bcf5edeb0e0a05b521d4abed4f8e507568bf59587f5442634d477c4

Pith citing papers

Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.701496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.701496Z digest=sha256:674da4b11d9314d08f4f59fa6f7910dc4da21c4c962db71adb170f8903739b9a

Observation 84a226cc-6e86-4bcb-af97-65ad555a9fe3 · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.741205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:eb86dcf3e3fea1948e2fb1b2085b89b67193dcde8b316f63417ffd6e7302fc77

Observation e0113f25-a5ff-4824-a94c-1ed1365b7eda · inbound

A Survey on LLM-based Conversational User Simulation cites this paper.

A Survey on LLM-based Conversational User Simulation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:11:12.091734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:21:09.118243Z digest=sha256:9ae5247685e0e6b7f57457010cbb5a42b835c5eddf73efe81db1db5307ca822a

Observation 18109681-da7c-4dc8-9741-66ce23d3569d · inbound

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation cites this paper.

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.182424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T12:54:56.401350Z digest=sha256:6d70f1e789b8bf4644018f8a8f95022fed0312fdba59a79a96f106314895c8a1

Observation 7364afbf-13de-463a-8464-87a822720ec4 · inbound

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking cites this paper.

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:52:52.347650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T19:51:45.657658Z digest=sha256:31ad4c0c9a1d9af2228e0b697afe793dd31ff9e896a9245290c0a61ea1ec3ec6

Observation bfa57893-7f3d-4d52-96d5-3d73151bf6d4 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.996301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:f5870afe143e8fe280908a5517a0549cb8522744634f4b9f30158696ab681cb5

Observation b7a42867-3b8e-4f87-b3a1-afb6ee7d0ab8 · inbound

Benchmarking the Personalization Capabilities of Large Language Models cites this paper.

Benchmarking the Personalization Capabilities of Large Language Models A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:09.892437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:17:09.892437Z digest=sha256:993f30f842079b7f66526840d1e36d2d2b7450dd92130fbc8661179666fdad09