Pith. sign in

Paper Citation Record · LEDGER

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.05246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05246 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:17:53.524604Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbe5b08b-a1a0-4570-84a2-5ea7da5335df · outbound

This paper cites Learning to reason for multi-step retrieval of personal context in personalized question answering.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Learning to reason for multi-step retrieval of personal context in personalized question answering

Reference 1

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:54.748244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.272051Z digest=sha256:ddefad42ad3c7b2ee4267e9f8779f0b3e5d20ff42e368e7af50eaaeb13e9f410

Observation 79ad8bfc-f4b9-429c-8877-4842ba5b5e35 · outbound

This paper cites Large language models empowered personalized web agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Large language models empowered personalized web agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.277727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.277727Z digest=sha256:09fa895b88a6ae89f3543344d60a72844b0c837fa7e6225b78370d9c6b6d5122

Observation 6b6bc56a-3184-44fe-a44e-f73234dc6201 · outbound

This paper cites Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.283149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.283149Z digest=sha256:cce8c260b07a190ae0f7c578e8dddf19e7e995b13500ab52def6ca5b8762e15e

Observation 05ac2751-17c0-4bf0-90b6-984eefa4c23c · outbound

This paper cites KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.289378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.289378Z digest=sha256:e6a6f17eb9090a8d8eaef28e6ee60fb2dfe8e32eff9de13492c973369f745fc3

Observation 647d158d-7195-422d-87bc-e353944c7d53 · outbound

This paper cites POPI: Personalizing LLMs via Optimized Natural Language Preference Inference.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs POPI: Personalizing LLMs via Optimized Natural Language Preference Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.294909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.294909Z digest=sha256:38fc0f74b84947acd066b68f4b3ce615c1f511db221a2cc53bf338b15915c226

Observation 5fa22ae1-bd35-49d5-a584-db57a9fdf609 · outbound

This paper cites Lifebench: A benchmark for long-horizon multi-source memory, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lifebench: A benchmark for long-horizon multi-source memory, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.301349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.301349Z digest=sha256:e056245c93057e9e96da38d840d912f2b5ff7035a236c18d4950d9dbd0cea489

Observation d566f43c-52a3-489b-b7d7-a44398374f6c · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.306961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.306961Z digest=sha256:c6e8c587ce5d50197389054128a0cf9cf59cfe00074ac46327b6df6e23ab79ed

Observation 1f13e3a0-abd7-4c66-8078-1ab509693913 · outbound

This paper cites Lifesim: Long-horizon user life simulator for personalized assistant evaluation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lifesim: Long-horizon user life simulator for personalized assistant evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.125871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.312072Z digest=sha256:21bdab7412c661eb2c66e86954c18a98aafd61a1629fd0fc456fdf2c86f03ccd

Observation 015c68cd-1c79-4cd7-87f4-2dc8ef03ff31 · outbound

This paper cites A survey on personalized alignment—the missing piece for large language models in real-world applications.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs A survey on personalized alignment—the missing piece for large language models in real-world applications

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.104016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.316730Z digest=sha256:780e1e5453007c9ea70809e42b356831ff1ff333bdfd8c63bc4a0a9f23c972a3

Observation 2f3e4e44-ae26-4184-bf7f-b33dbd09a03a · outbound

This paper cites Towards realistic personalization: Evaluating long-horizon preference following in personalized user-llm interactions.arXiv preprint arXiv:2603.04191, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Towards realistic personalization: Evaluating long-horizon preference following in personalized user-llm interactions.arXiv preprint arXiv:2603.04191, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.321693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.321693Z digest=sha256:23229029ab74f7f37cb8b0d901a6c75a10e39ed3d90fae7484a4da0439e58714

Observation 5a833f00-6104-4aff-8872-1a955379b6ea · outbound

This paper cites Computing inter-rater reliability and its variance in the presence of high agreement.British Journal of Mathematical and Statistical Psychology, 61(1):29–48, 2008.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Computing inter-rater reliability and its variance in the presence of high agreement.British Journal of Mathematical and Statistical Psychology, 61(1):29–48, 2008

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.327020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.327020Z digest=sha256:e4e1d990c44daeaf54dc699c7f980d76751916841e6a992b3fdce02d2b8433a7

Observation 4e65dcc8-0313-48e6-ba2f-143f54d7cea1 · outbound

This paper cites Rap: Retrieval-augmented personal- ization for multimodal large language models.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Rap: Retrieval-augmented personal- ization for multimodal large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.074174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.331856Z digest=sha256:ba8328c7b80771b3da8ea3fcd08f5ae20218f6d16f4963927339c75499e49035

Observation 56da7c87-ee3c-4df9-b752-e35cd45242dc · outbound

This paper cites Asking the Right Questions: Improving Reasoning with Generated Stepping Stones.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Asking the Right Questions: Improving Reasoning with Generated Stepping Stones

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.336566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.336566Z digest=sha256:00ad8e02c1ef2129e2d6ff3a7a96f2ef6fe7174487b3ac8e820bd941bb1a774b

Observation 587f4df2-c4fc-4f95-8a4f-e51722811664 · outbound

This paper cites Op-bench: Benchmarking over-personalization for memory-augmented personalized conversational agents.arXiv preprint arXiv:2601.13722, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Op-bench: Benchmarking over-personalization for memory-augmented personalized conversational agents.arXiv preprint arXiv:2601.13722, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.342011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.342011Z digest=sha256:a6177e190a107d6602ad942a56ad2711665df5e902601e1e74e3ecb126de0146

Observation 26d5f4ea-1045-4c42-a35a-452f625912d9 · outbound

This paper cites Mem-pal: Towards memory-based personalized dialogue assistants for long-term user-agent interaction.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Mem-pal: Towards memory-based personalized dialogue assistants for long-term user-agent interaction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.056133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.346834Z digest=sha256:44be0d6a527528e55de9e39be663c2550e65bedf0cf4bb76019eb3f6467ca2cc

Observation fc44df1e-9e30-44e3-8dea-6c4ea85bddbd · outbound

This paper cites Taylor, and Dan Roth.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Taylor, and Dan Roth

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.352101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.352101Z digest=sha256:866857990e6811a5728cfb2e75b5d7eb35db53db711b6749e8afb3cdb75f76bd

Observation 09fed038-012f-4d9a-a4e7-d0b7da1e84b5 · outbound

This paper cites Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.356534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.356534Z digest=sha256:fd1002ba0a3fe7ceb1b84c65d86493eadd2fe0caf40ac1abb295edec896bcc08

Observation 1761f197-9cf5-44b7-9579-ff4770ea77d5 · outbound

This paper cites Humanllm: Towards personalized understanding and simulation of human nature.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Humanllm: Towards personalized understanding and simulation of human nature

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:54.163090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.361281Z digest=sha256:7e78c8b8a96ba6ab050fae21654610b1b9756deb8277dd267ede7444be119d4f

Observation 3df11aa5-64a0-4e50-9c98-4ab4a69a979f · outbound

This paper cites Retrieval-augmented genera- tion for knowledge-intensive nlp tasks.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Retrieval-augmented genera- tion for knowledge-intensive nlp tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.039109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.366074Z digest=sha256:1a7a55476c2efd9bdc82c46232824ff902c7af93f0e54b51b9fe5928ffe8fb24

Observation 0ec4ee73-f0bd-48d8-b716-81e5dd09304c · outbound

This paper cites Can llm agents simulate multi-turn human behavior? evidence from real online customer behavior data.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Can llm agents simulate multi-turn human behavior? evidence from real online customer behavior data

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.021118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.370929Z digest=sha256:3490d5cda6868a64f79555e29b0c39137b83e7cebaf21d06161779d473d41d76

Observation 1f0c53b9-6e91-4f50-94c1-229004c058e5 · outbound

This paper cites Exploring the potential of LLMs as person- alized assistants: Dataset, evaluation, and analysis.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Exploring the potential of LLMs as person- alized assistants: Dataset, evaluation, and analysis

Reference 21

Resolution
verified exact
doi, observed 2026-08-08T17:17:53.595425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.375836Z digest=sha256:7de31649a9c95cb505d77ebd0578730b468455ed01dc0b87f4c1a4852bdc2958

Observation d126780d-3f88-4567-a0a9-0fda6c6b59fc · outbound

This paper cites Privacybench: A conversational benchmark for evaluating privacy in personalized ai, 2025.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Privacybench: A conversational benchmark for evaluating privacy in personalized ai, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.380675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.380675Z digest=sha256:665ba89ba4c0553c1e7ac17a99642088d9380923786c5866973c688d1731962e

Observation ad1f9ba3-504a-4ab0-8a64-429516c6ef1e · outbound

This paper cites PersonaVLM: Long-Term Personalized Multimodal LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaVLM: Long-Term Personalized Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.385274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.385274Z digest=sha256:303900dca74a36101af0ec90ff970fcac39f9d27500a539ed399d5fb9e4fe64d

Observation cb7d46fd-ed79-4425-ae21-9c923116454c · outbound

This paper cites On Memory Construction and Retrieval for Personalized Conversational Agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs On Memory Construction and Retrieval for Personalized Conversational Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.390227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.390227Z digest=sha256:101db1a3cc434a38b607d9d8cb0ad672aaaffe412bd962aee833f8dd0a340c6e

Observation a2e0a7b9-8a7c-4286-a73a-a35006f2af53 · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.395160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.395160Z digest=sha256:5d0c0bf562e54a743fede4176d92fd168cf32c055d7203ad3503c303937d23c6

Observation c4df0f79-a0e7-40ed-aa7c-a6b0fc3768e0 · outbound

This paper cites Lamp-qa: A benchmark for personalized long-form question answering.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lamp-qa: A benchmark for personalized long-form question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.003847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.400822Z digest=sha256:c0571cafff96145095625ba54ceb4f369b8b04c357c0be128f58e5e2029093b5

Observation b7cb8012-50dc-419e-bf7d-a489e1e3d0e5 · outbound

This paper cites Optimization methods for personalizing large language models through retrieval augmentation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Optimization methods for personalizing large language models through retrieval augmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.405758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.405758Z digest=sha256:55383b4aa70a75ca114a9be306a8643e76a885844877598a13a2cf67c76fd00e

Observation 636e4262-22b5-469c-be69-20b8decfc0c2 · outbound

This paper cites Lamp: When large language models meet personalization.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lamp: When large language models meet personalization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.987973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.410357Z digest=sha256:73825e4dae6d6c04bf79ddf3d4282aa64d7057c99d7571ae05bf7874f6d14d0d

Observation 63cdc21d-4945-41e9-b83a-2d50e16f610b · outbound

This paper cites PersonaBench: Evaluating AI models on understanding personal information through accessing (synthetic) private user data.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaBench: Evaluating AI models on understanding personal information through accessing (synthetic) private user data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.414867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.414867Z digest=sha256:f3651ba63f5b9624f652bb57db56ddc8ce2bbb063b3d99d2f49e429b794d5ff8

Observation ad19aebd-878a-49c0-a267-9cf724497249 · outbound

This paper cites Democratizing large language models via personalized parameter-efficient fine-tuning.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Democratizing large language models via personalized parameter-efficient fine-tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.961165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.419312Z digest=sha256:564c8df9db6d0d37d502f629e01f4d31cf0c95591413be9215108c267be8b70b

Observation 5e34d33a-f1a6-4e94-88bc-38449f3a8511 · outbound

This paper cites In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.945799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.424373Z digest=sha256:a52d33185f8895989120aad74accc06c8ee12eae339c7cd8b54e77ae5935eb11

Observation cba440d3-9b69-45e2-a3f3-0aa57e8bb554 · outbound

This paper cites PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.434087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.434087Z digest=sha256:99f5592f0a58a990d6325d4408f90442485b05cdefde4416ef5195cd08f7a365

Observation d24e8727-09ee-4441-ab20-6213f38d17de · outbound

This paper cites OPeRA: A dataset of observation, persona, rationale, and action for evaluating LLMs on human online shopping behavior simulation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs OPeRA: A dataset of observation, persona, rationale, and action for evaluating LLMs on human online shopping behavior simulation

Reference 33

Resolution
verified exact
doi, observed 2026-08-08T17:17:54.930558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.438950Z digest=sha256:230c4e3a9adac71bafa2da94ae27cc029362ba224eacb73b64431e291668e801

Observation d7cabc12-5fd3-4ad5-9e9c-693f2d03cd75 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.443925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.443925Z digest=sha256:61bfded05cf75978e2c620422d394fcb49601c77fa1a0242f2f3c96eb7331067

Observation 1636cc8f-8755-4d64-a7a7-2d33cc89459b · outbound

This paper cites DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.448806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.448806Z digest=sha256:6f507ff6599dee5ecd6a9eaecc1e109eb1f50ef417b329fbdb99b0756119c918

Observation 294eab45-e44f-4b36-b70a-debc93e08f27 · outbound

This paper cites Lauvrak, Jon Atle Gulla, and Heri Ramampiaro.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lauvrak, Jon Atle Gulla, and Heri Ramampiaro

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.913126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.453682Z digest=sha256:ead976d50721d4090f9e37a550efe122a1844b9967b753365cc1e01f451a8062

Observation d87058c1-ad0c-498c-956d-3ab710d15acf · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.459192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.459192Z digest=sha256:c5bf10c57189ba4ed9a9098f7ab3973d11259270e0268b8d937953dbbbfc692a

Observation 19ce0ba7-0d2c-4154-8a0b-359a578a26b8 · outbound

This paper cites Promax: Exploring the potential of llm-derived profiles with distribution shaping for recommender systems.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Promax: Exploring the potential of llm-derived profiles with distribution shaping for recommender systems

Reference 38

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:53.739143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.464402Z digest=sha256:e35f3de62c250503b9881f8b0ccef07c364e6610c66a8f301767e66e4a4b91d6

Observation adffac51-bba2-4278-aec2-65d48b3e81bd · outbound

This paper cites Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.469368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.469368Z digest=sha256:e477277c8940ab6843f86c9f71373b7f39b6f938ee634916f1c298c79a265cb5

Observation d0a33fd7-19bc-4d37-b060-191a6d177e0e · outbound

This paper cites Cohen, and Emine Yilmaz.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Cohen, and Emine Yilmaz

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.474220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.474220Z digest=sha256:993338152e5b084fdf0bc63699ff503cd25396da303464c353018078037b5f33

Observation 3bef6ec4-4c44-4604-9663-e53fa6a84e70 · outbound

This paper cites Cantonese cuisine.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Cantonese cuisine

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.479258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.479258Z digest=sha256:190d14701d3145984875f886bbc1f9179ff1be71670f86d5882dcdd364d371c6

Observation 46159ad9-2549-44f9-819d-091c4950fd99 · outbound

This paper cites You tend toward budget-conscious choices.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs You tend toward budget-conscious choices

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.896381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.485499Z digest=sha256:92fd37acf6f3b026991104161874556538b14ada5217144a3b12c8cc0b22598b

Observation cbacb1d9-02bc-4d7d-911f-7b2fd1fbb502 · outbound

This paper cites Note that cross-domain references serving the response are legitimate; offense occurs only when references are purely demonstrative.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Note that cross-domain references serving the response are legitimate; offense occurs only when references are purely demonstrative

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.880274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.490402Z digest=sha256:409eec13fe190bd87475e356660da6269d2d337fb97f8b4f1f77193c7cda415f

Observation 91dd12f4-3f53-478f-99a7-0b60caec758d · outbound

This paper cites educate" the user, such as.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs educate" the user, such as

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.863931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.495608Z digest=sha256:37a0f60c4509a94e284f4dba02909e322b4c011a49065101a608ecc585476cc6

Observation 8d064032-1750-4f6c-9512-85b3a0f62214 · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.848030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.500517Z digest=sha256:4cfe0b02c89ccf147c8f1a1c54d3f1ae6295562a04f6124a3fc93ed1b3c030eb

Observation ad1790ac-5891-4b0f-ad40-1fab75938661 · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.830944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.505050Z digest=sha256:fecbd16166f51d1524fa8eec703e17f4efb5739d15492fe4b1756608b2a41c1f

Observation 177d19fa-f24d-4b90-bd51-6c8034977286 · outbound

This paper cites retrieval.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.815062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.509773Z digest=sha256:32e5f226db04bf6b60eddf79e04e571c4f28a210b94a28624261d0aedf839e93

Observation c704c8f8-0641-4b77-8a85-e02176aefcae · outbound

This paper cites swap test.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs swap test

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.799387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.515287Z digest=sha256:e1d1471c31d058fdf61a002d6420e747a995f44c6e1518e23d7a7c47a024a92a

Observation f44d3304-2982-4065-862d-6d738a01164c · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.781914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.519873Z digest=sha256:76a123c750cc96c61274021d0903dcc9cb9856e4b00abca9505429be3693131d

Observation 7f5c301d-97ed-471f-a7b4-055c5b9993f3 · outbound

This paper cites retrieval.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.765834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T17:17:53.524604Z digest=sha256:9a875f32b949f786053836af935f5ad6b1f41d599e3908e0443515b99cbfd411

Observation 0e0385c0-3997-4f02-8d7f-73c7f45ff40a · outbound

This paper cites ISBN 979-8-89176-251-0.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs ISBN 979-8-89176-251-0

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.429103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.429103Z digest=sha256:dc608a3a7e9349bf3d8b87e5ebb2816ddd00de814524e65d716900d96093d95e

Pith citing papers

No inbound Pith citation observations are available.