Pith. sign in

Paper Citation Record · LEDGER

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2505.14106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14106 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:23.733585Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.701496Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:34.994786Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9861fdb-99de-45c2-8df0-515c31a77adb · outbound

This paper cites Conversational Health Agents: A Personalized LLM-Powered Agent Framework.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Conversational Health Agents: A Personalized LLM-Powered Agent Framework

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:18.519445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:18.519445Z digest=sha256:42927d868529db3c606bfea84b3a531233d58d2575b5816e7c24e7642d7607c4

Observation e00bf5c9-bc0c-47b8-b2f7-cd2eb10dc9d3 · outbound

This paper cites Persobench: Benchmarking personalized response generation in large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persobench: Benchmarking personalized response generation in large language models, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.054834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:18.625291Z digest=sha256:41e330b32223284753536b7036c6cbe946df0a2738fa7d8815cf04d44abbeddc

Observation 03167942-a7e7-48c2-bf1e-e6ae533ea020 · outbound

This paper cites Claude 3.5 sonnet.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.909324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:18.729628Z digest=sha256:8f3750058aa9e926e9dfb753ee204e2b84a1b41b75182dc322ccd85efe872a0a

Observation 237bf522-348b-4651-a528-b532151c0119 · outbound

This paper cites A Little Human Data Goes A Long Way.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Little Human Data Goes A Long Way

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:25.276669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:18.993776Z digest=sha256:f53c40b3f5584ee285a57b4163278126ed8a799140e5d0b1f1c782190c7110bc

Observation c3e01726-ba8e-4cf9-b027-26cfc2176d36 · outbound

This paper cites Personalized Graph-Based Retrieval for Large Language Models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Graph-Based Retrieval for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.113061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.113061Z digest=sha256:74dbdd928e8c4b2096499abceab5e8a76d7e4f8369157da49107e2f30aabc7e2

Observation 03232c27-8236-4033-b7d8-f3d2f54d19af · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.244719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.244719Z digest=sha256:6b33e9b42761f3800add7b19cbc9f772cd68dca1e3eaf6f9a63fbe7a15a86e57

Observation a903f797-26e0-48dc-b27f-1e35298c32d9 · outbound

This paper cites LoRe: Personalizing LLMs via Low-Rank Reward Modeling.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations LoRe: Personalizing LLMs via Low-Rank Reward Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.367275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.367275Z digest=sha256:898a7b67325e7cc2a3202c31f21b50285be132ad0991f4756d87818532aeee29

Observation a8722fdf-21c2-468e-b36a-a02ebb7c345c · outbound

This paper cites Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.490084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:19.513223Z digest=sha256:bafd4eac286a8125baea2e8ee122dcf72f0aa2d315f237368c43fde5becd92e5

Observation 74b080f6-b2ed-42c0-a046-72636d8e93b4 · outbound

This paper cites Beyond prompts: Dy- namic conversational benchmarking of large language models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Beyond prompts: Dy- namic conversational benchmarking of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.268054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:19.655393Z digest=sha256:2d7e6bb43c4c13dad6a28ed515cf318fef6c33a7ec5979f089ad664cf4af02a2

Observation f3e6934d-15cb-4155-9c60-ad8fd8a2b49f · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.075713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:19.807578Z digest=sha256:076b00b79bb389c60bb23cbfd06df51dcb3ff315bf15b1b5050c0fde63f5b425

Observation 2ef064e0-e58f-4565-9410-fc52a2f1045c · outbound

This paper cites When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.894818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.894818Z digest=sha256:899c4392f7260acc20bc00e5b2c8e6ca8ec8b291372f4114b3af9b404ea74692

Observation 3cc7d673-0b61-4547-aa25-30019abbd5a8 · outbound

This paper cites REALM: A Dataset of Real-World LLM Use Cases.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations REALM: A Dataset of Real-World LLM Use Cases

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.811302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.033370Z digest=sha256:28e9fe109cb7c6d83c124851cbcee7054961b7ac3354566a5428f5e67fe7f491

Observation 4e070871-557d-43a0-adea-d59997ee4ab3 · outbound

This paper cites The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.909265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.144923Z digest=sha256:df65f977a481cc92b02ddd89cf1a7509edd4875cf956702fcecb4ceb9b85f85b

Observation b06db013-27d0-4816-9cb6-68ce7003e4d0 · outbound

This paper cites The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.761772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.324986Z digest=sha256:968d86681cec4ae63d1ba417ab74ca56f382910390e929d0225bfefbec8a0943

Observation 88126155-e729-4f8f-8fc5-f140559fa286 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.595175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.465431Z digest=sha256:6bf2de4fc3d13b97c4fefa3211a029052f3638913722b5b1d13f9ec06e1126a3

Observation c48a18f2-6a9e-40d5-b11d-dd9a31c50a36 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations RedCaps: web-curated image-text data created by the people, for the people

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.580260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.580260Z digest=sha256:28522b02e74bf53e2a6b87e9d2bae36ae0ff32ed25d4afdfca2b8730ea1f21fc

Observation 9a06c4d6-7fac-4a12-a0aa-9fc9e604ac27 · outbound

This paper cites The llama 3 herd of models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The llama 3 herd of models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.662819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.662819Z digest=sha256:4f451a0792e37613ef9df7e8a13f932d16f87c92f8b1187b2194eaea74525ce5

Observation 850ceb35-110c-4dc9-9b9a-a27a373e84d2 · outbound

This paper cites Ruddit: Norms of Offensiveness for English Reddit Comments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ruddit: Norms of Offensiveness for English Reddit Comments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.552148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.731343Z digest=sha256:83f7facf5e8e0d9ffd52d34aad85e23060b326d9bc34a5859a368371c08560ca

Observation 194d8e5f-4214-4e7e-bfc7-0a4e1ee518eb · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.382963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.867591Z digest=sha256:d60e89531e3be452f115a080d1a1d05a83978b0fea20f0ded0aee56d570b28f8

Observation e4f4b3bd-d184-45bd-9946-d970bf281d4c · outbound

This paper cites Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.165335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:20.984363Z digest=sha256:4ce07649d52ea59044d975c3080bb14243538a9f00a3c0116ecf500bcfd523ec

Observation 570f0968-3ae8-4466-8fbe-45f11d61adc4 · outbound

This paper cites Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.936888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.117526Z digest=sha256:021b65f352f2b0cea878a5a342c9d3dbddc4c0724f4568a735680ca9f500a23a

Observation b6daae78-9455-4ebb-ac01-1c48db55eb46 · outbound

This paper cites A framework for building adaptive intelligent virtual assistants.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A framework for building adaptive intelligent virtual assistants

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.672196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.243210Z digest=sha256:6a5fb4d9bd80dde792de4d07091c76c7383add28697082ae343ed5f7c420c4ad

Observation 77ecc1ea-0684-4054-b46c-9b666cf43b3c · outbound

This paper cites Teach LLMs to Personalize -- An Approach inspired by Writing Education.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Teach LLMs to Personalize -- An Approach inspired by Writing Education

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.349336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.349336Z digest=sha256:7374e10382a483688009168e14661774a42605b27ac5b80bf700d595aae85e88

Observation 1945642b-fd5f-49bf-a9bd-77d49420dab2 · outbound

This paper cites Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.437751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.487031Z digest=sha256:862659a79b06b53fdb384a29402d6b6fa9150b924fae3c695ec5302568a2066a

Observation 62133eab-494d-409c-9d3b-59c8a7b733aa · outbound

This paper cites Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.240242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.584797Z digest=sha256:3ce5c00d459db0eff0a6812c1b55b78cf1b777b93818ec6367844952f0f4c4b3

Observation 955fdb45-569b-4517-b7be-6a5169debecc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rouge: A package for automatic evaluation of summaries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.651056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.651056Z digest=sha256:9fd4dbae98b9bd191933eae02814cffac4b7e8b9a84f58e8ac1b400ec72c11b2

Observation 646fc09e-b860-47e2-a683-92459fac4bb5 · outbound

This paper cites Persona-sq: A personalized suggested question generation framework for real-world documents, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persona-sq: A personalized suggested question generation framework for real-world documents, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.971000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.760648Z digest=sha256:d75fba5b33a7b3b840c95b57cc785eb970af4a23b2a62e793987c61f0e25d2d4

Observation 9533affa-216e-4632-8ac6-64befd694cc9 · outbound

This paper cites Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.829533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.906459Z digest=sha256:ed6499f2f12ce5d156381662632942d8ed8fae350f1f6e3e72f53cf15ba22402

Observation a054c5ee-571e-44a1-bd79-40f20b53da15 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Gpt-4o mini: advancing cost-efficient intelligence

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.643229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:21.975869Z digest=sha256:17aa9b61b0fa1fe5955630b2030cff5b121b7a279322829118ce64fc66a9a4b8

Observation 6ac76a55-64bc-42c0-96ca-c0cfd0f0339c · outbound

This paper cites Introducing gpt-4.1 in the api.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Introducing gpt-4.1 in the api

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.480886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.043964Z digest=sha256:b8579f9a74055bc06601842b251dee42152ab956b1bbba5a0508ce3aa3ce80bd

Observation 9d69a277-3a48-4798-855c-a90eab1feb4e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Bleu: a method for automatic evaluation of machine translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.172514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.172514Z digest=sha256:e0c278fc3603b9427a71673e19db6433be2d76184fb82985d5ed4eddeefa3229

Observation 65e016f6-564b-4426-8085-d6a390ff8535 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:31.250804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.264169Z digest=sha256:b5e3bb46cfa358127d69972032850c58b36c3d9834253a2dc381badfa4c94f32

Observation 0fed4877-ccbd-4a02-9aee-37490e77aec9 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:30.988064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.394518Z digest=sha256:8b0c158549ba11cd858bcbb9269916afd49d1bc5da988ca2d7b61fa2b9986927

Observation e44488dd-ecf2-4aaa-83fb-0f4bdac730a5 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.486570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.486570Z digest=sha256:c5273e51de0c3fb930e03a357e7bb916e39d0fedda1881665d5f43ba80aabe18

Observation e08ac11f-6f2a-4ef2-bcd4-b722c05ab059 · outbound

This paper cites Lamp: When large language models meet personalization, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Lamp: When large language models meet personalization, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.780537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.581458Z digest=sha256:67f11a6c048e414ba39bc2cf2a5591105a0d6d90d06b4c3bf443cb2e5a68db23

Observation 639f9236-0fca-4285-9c56-1765c052d4e5 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.649086Z digest=sha256:6548d6fa278f58e882904301e7ea69cf452185a4bd5ee6c4aabdf9fe76dc52d2

Observation dbf2d392-76b7-4210-bacf-45ab18e9a5db · outbound

This paper cites Democra- tizing large language models via personalized parameter-efficient fine-tuning.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Democra- tizing large language models via personalized parameter-efficient fine-tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.623584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.762396Z digest=sha256:955a774f48690e608698a9270a17c63eca2e69c744665b66828276e1aafd6659

Observation 2d87413e-c6d6-43ab-9c83-59a6cddc2c4e · outbound

This paper cites An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.404821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:22.887964Z digest=sha256:fd3a63ca0000c6103ace2cd88ef1d117510ea70c2f7d627ed86ed5202b360292

Observation 6b15c0d2-b97b-46a3-abd4-7495a23729a5 · outbound

This paper cites Position: Will we run out of data? limits of llm scaling based on human-generated data.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Position: Will we run out of data? limits of llm scaling based on human-generated data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.943347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.943347Z digest=sha256:de219cdd104536446910168f230dea5baacf2c726c7f30411832ae7a2fb59c5e

Observation 2d03cefe-d484-4bf3-8557-0007abc5c064 · outbound

This paper cites Personalized Multimodal Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Multimodal Large Language Models: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.001354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.001354Z digest=sha256:26cfbc6d60ca341c28326b7d389b9507f6ae02d471c2a0054b98e58f434caa24

Observation 0c64fbc1-cfc5-4224-ba00-99c35c647e1c · outbound

This paper cites A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.118282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.118282Z digest=sha256:34c7434b0e58a8298c806debcafe9c22ab8ad4c57c63d1ccc3fb1b760faae5a8

Observation f5e06eb9-0d06-48ae-bf85-deb4ba1386b5 · outbound

This paper cites xdial-eval: A multilingual open-domain dialogue evaluation benchmark.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations xdial-eval: A multilingual open-domain dialogue evaluation benchmark

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.185319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:23.196806Z digest=sha256:f844b32a42537f6364c7b6d0ccfd8ec36c1d5c69e64ae2e5d9b33d09ea3f907d

Observation ca8f1e77-8842-440d-a565-21323465c189 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:25.986776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:23.314664Z digest=sha256:69b4b16ae10c7d7091ce1a5bf761cb7ad88418219733707ca566a872814cc80c

Observation e23151f7-29ff-40fe-9dbb-d849f8ae67e2 · outbound

This paper cites Personalization of Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalization of Large Language Models: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.443293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.443293Z digest=sha256:ecdc0f6062fad9f864afd3867d7612ade3ef950c9ab708f0fffb7f2ef6543cfe

Observation a9e526ef-fff8-4464-832b-32f5844900cb · outbound

This paper cites DiQAD: A benchmark dataset for open-domain dialogue quality assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations DiQAD: A benchmark dataset for open-domain dialogue quality assessment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.739429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:23.527996Z digest=sha256:9f9da925f2313b06b787c819d89f2485c9d15c4c0e5a9cae20a53a57ee3d27ef

Observation 791e650c-e121-426e-87d2-3373f14446e5 · outbound

This paper cites Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.537365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:23.662353Z digest=sha256:5d6bc691453e7e9e9d9efd9ac25eeeb0d591cbfa42d5e975c3f4f359340fe1e1

Observation 7c921583-99e9-4d70-8560-649b95daa7cc · outbound

This paper cites Best Response.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Best Response

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:43:24.106045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:23.733585Z digest=sha256:862faabdd45d5cef2d77fe91b552de205a116189b678d9f2362e82e83049fcce

Observation 3c5c23a3-3e7f-4e88-9235-16ffc7f9185b · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:43:34.736747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:43:18.862204Z digest=sha256:4b94d6a2eff42dc302a3fe39799e982e2cd0b47faac8d1eba9a4a794eabb400b

Pith citing papers

Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.701496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.701496Z digest=sha256:9f9b16ef3901d2aae6b868611a3dddd9308c8b6f9daf9365f36948d203d70e0d

Observation 84a226cc-6e86-4bcb-af97-65ad555a9fe3 · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.741205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:90a8763bd0f93caf49e17dfc29de6b40bc605a60beac274c0306be61040ca01c

Observation e0113f25-a5ff-4824-a94c-1ed1365b7eda · inbound

A Survey on LLM-based Conversational User Simulation cites this paper.

A Survey on LLM-based Conversational User Simulation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:11:12.091734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T03:21:09.118243Z digest=sha256:1f6fa3809016652137858e0e2863a41f06d9ead3c6aaebc4ea345885a46c097b

Observation 18109681-da7c-4dc8-9741-66ce23d3569d · inbound

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation cites this paper.

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.182424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T12:54:56.401350Z digest=sha256:b62f688718d712fb89a08ff77d73f7d7eda67ad360976d77211070d51b6adaec

Observation 7364afbf-13de-463a-8464-87a822720ec4 · inbound

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking cites this paper.

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:52:52.347650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T19:51:45.657658Z digest=sha256:d861e46063936d439de34f2698a94d289303d1afc6b4b59b427b7f223e7280ec

Observation bfa57893-7f3d-4d52-96d5-3d73151bf6d4 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.996301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:7aac66384c1e5ece8d7d82b2e81c9e6d9dc5c2b714147ce4fb629a629fc400d4

Observation b7a42867-3b8e-4f87-b3a1-afb6ee7d0ab8 · inbound

Benchmarking the Personalization Capabilities of Large Language Models cites this paper.

Benchmarking the Personalization Capabilities of Large Language Models A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:09.892437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:17:09.892437Z digest=sha256:18063f2349414134f579e45f4f722100ba0fb18b1a57bf33d41379510db464d5