Pith. sign in

REVIEW 5 major objections 5 minor 95 references

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Behavioral traces alone can yield a persona that matches or beats interview-built personas on next-action prediction.

desk verdict The IToM pipeline is a well-described, honest contribution to LLM-based user modeling, but its headline claim of beating ground-truth personas rests on a format-confounded comparison the authors themselves concede. read the letter →

arxiv 2608.11354 v1 pith:W4FOM4OJ submitted 2026-08-11 cs.AI

classification cs.AI
keywords InverseTheoryofMindpersonasynthesisrecommendersystemsabductiveinferenceLLMreasoningBigFivepersonalitygenerativeUIusermodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a user's 'why' — the beliefs and decision style behind a click — can be recovered from ordinary interaction traces without interviews or surveys. It proposes an Inverse Theory of Mind pipeline that reconstructs what a user saw and chose, generates natural-language counterfactual belief statements, and composes those beliefs into a portable persona. On a 46-user online shopping dataset with interview and personality ground truth, the inferred personas match or exceed interview-based personas at predicting the next action, improve Big Five prediction over population norms, and align with self-reported shopping attitudes. The payoff for a sympathetic reader is a user model that is interpretable, auditable, and transferable to new interface modalities such as generative layouts or spatial interfaces, where click histories alone are not enough.

What carries the argument

The load-bearing object is the reconstructed percept: the decision context in which an action occurred, rendered as the main choices, sub choices, other available options, and the user's actual choice in natural language. Each inferred belief must cite a chosen option and a rejected alternative from that percept, which grounds the counterfactual reasoning and is intended to dilute hallucination. The pipeline's second load-bearing mechanism is multi-hypothesis persona synthesis: it generates several distinct personality interpretations of the same evidence, scores each by quality and per-dimension evidence confidence, and combines them with an adaptive population prior, so weak evidence regresses toward published norms instead of an overconfident stereotype. Belief optimization with maximal marginal relevance selects a compact, diverse belief set in between.

What would settle it

If a user's retrospective think-aloud rationale for a specific action disagrees with the LLM's inferred belief for that action at chance level, or if feeding the pipeline only the chosen action without the reconstructed alternative set leaves downstream accuracy unchanged, then the inverse-reasoning mechanism is not what carries the result.

Watch

Extended reading notes

Core claim

The central claim is that persona-level understanding can be inferred, not elicited: a five-stage pipeline (Summarize, Perceive, Infer, Optimize, Synthesize) converts raw action events into a structured persona, and that persona, on one frontier LLM backend, raises exact next-action generation accuracy from 17.06% with the ground-truth interview persona to 25.69% with the inferred persona, a 50.6% relative gain. The same pipeline lowers Big Five mean absolute error to 0.762, below the 0.836 of predicting population norms, and reaches 76.6% normalized accuracy on shopping-attitude items. A further claim is that single-interpretation LLM inference from behavioral personas is systematically biased on this cohort, producing predictions anti-correlated with ground truth, while aggregating several diverse persona hypotheses with confidence weighting reverses the correlation sign. The paper reads the action-prediction advantage as evidence of a representational-format effect rather than proof that inferred personas understand users better than interviews, and it grounds that reading in the episodic, belief-anchored format of the inferred personas.

Load-bearing premise

The load-bearing premise is that the LLM's counterfactual belief statements faithfully reflect the user's actual decision process rather than plausible-sounding stereotypes, and the paper reports no direct test of that faithfulness.

Editorial extensions

If this is right

  • Interview- and survey-based persona construction becomes optional for many downstream tasks, since behaviorally inferred personas carry enough signal for next-action, attitude, and category prediction.
  • A single natural-language persona can be reused across tasks and interface modalities without retraining, so a persona built from screen browsing can steer layout and content in a spatial interface.
  • Multi-hypothesis aggregation provides a floor against LLM stereotyping: when behavioral evidence is weak, predictions regress to population priors instead of confident anti-correlated guesses.
  • Episode-grounded, belief-anchored personas are less likely to inject role-play noise into step-level action prediction than abstract trait summaries.
  • Held-out shopping category prediction reaches Hit@3 of 100% and NDCG@5 of 0.701, suggesting the persona generalizes to sessions it was not built from.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the pipeline's belief statements turn out to be largely stereotype-driven, the reported downstream gains could survive even when the inverse-reasoning story is false; a belief-level faithfulness test against users' own retrospective rationales would separate the two.
  • Beyond the paper: the action-generation margin may partly be a prompt-format artifact of episodic phrasing, as the paper itself notes; a format-controlled comparison that rewrites ground-truth personas into the same episodic style would isolate the user-modeling contribution.
  • Beyond the paper: the percept-reconstruction step should generalize to traces without page HTML, such as eye tracking in spatial interfaces, gesture logs, or typed queries, where a test of the same pipeline on a non-shopping domain is the natural next experiment.
  • Beyond the paper: if persona portability holds, recommender systems could adopt a compute-once, query-anywhere pattern, paying a one-time inference cost per user and then reusing the persona across every adaptive surface, changing the cost model of personalization.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The manuscript proposes an Inverse Theory of Mind (IToM) pipeline that reconstructs a user's decision context from raw interaction traces, uses an LLM to generate counterfactual belief statements, optimizes and deduplicates these beliefs, and synthesizes them into a structured natural-language persona through multi-hypothesis aggregation. The pipeline is evaluated on the OPeRA dataset across four tasks: next action prediction, held-out shopping category prediction, Big Five personality inference, and shopping attitude survey alignment. The authors report that inferred personas match or exceed ground-truth interview-based personas on next action prediction, that multi-hypothesis reasoning reverses a negative correlation between LLM-based personality estimates and ground truth, and that the resulting persona transfers to a VisionOS spatial banking prototype. The paper is clearly written and honestly acknowledges several limitations, but the central empirical claims rely on small samples and on at least one comparison that the authors themselves concede is format-confounded.

Significance. If the central claims held, the contribution would be significant: a behavioral-trace-to-persona pipeline that is interpretable, modality-agnostic, and reusable across tasks would address a real bottleneck in recommender systems and adaptive interfaces. The use of an external dataset with ground-truth interviews and personality assessments is a genuine strength, as are the direct comparisons against external published Big Five norms and the transparent reporting of LLM call costs. The multi-hypothesis aggregation idea is interesting and, if properly validated, could be a useful corrective to the known tendency of LLMs to produce stereotyped personality judgments. However, the current evidence is not yet at the level of the paper's headline claims: sample sizes are small (N=46, 27, 15), several comparisons lack uncertainty estimates, and the key RQ1 comparison is explicitly admitted to be confounded by prompt format. The paper is therefore best viewed as a promising proposal that first needs a format-controlled validation and more careful statistical reporting.

major comments (5)
  1. [§4.3, Table 1; §4.7] The paper's headline claim that inferred personas 'match or exceed ground-truth personas' is not established by RQ1. The GT persona is an abstract trait summary, whereas the inferred persona is an episode-grounded natural-language profile; Section 4.7 explicitly concedes that the +50.6% relative gain on Claude Opus 4.6 'may partly reflect prompt-format alignment with the action-prediction task' and that a format-controlled comparison is future work. Because Table 1 is the primary evidence for the abstract's central claim, the comparison needs a control that rewrites GT personas in the inferred-persona format, and an ablation that removes episodic content from inferred personas. The table should also report bootstrap confidence intervals or a significance test over the 15 users / 90 sessions / 985 instances, since the 25.69 vs. 17.06 difference has no stated uncertainty.
  2. [§4.5, Table 4 and Figure 1] The claimed benefit of multi-hypothesis aggregation for Big Five prediction is the sign reversal of the average Pearson correlation (−0.103 or −0.008 to +0.103). However, the aggregate bootstrap CI is reported as [−0.05, +0.26], which includes zero and negative values, and the only per-trait correlation approaching significance is Extraversion (r=+0.280, p=0.060) without any multiple-comparison correction. With N=46, this does not establish that the sign reversal is a stable population-level effect. I ask for per-trait confidence intervals, a bootstrap distribution of the aggregate correlation, and a statement of how many of the five traits have CIs excluding zero.
  3. [§3.8, Eq. (5)] The adaptive prior strength α_d and the confidence weights w_h are central to the reported personality and attitude results, but they are underspecified. The text says only that α_d 'scales inversely with mean evidence confidence,' with no formula, no default values, and no description of whether α_d or w_h were calibrated on the 46 evaluation users. If these weights were chosen after inspecting the results, then RQ3 and RQ4 are not held-out evaluations. Please specify the exact computation of α_d and w_h, report their values, and include a sensitivity analysis over a plausible range of α_d; alternatively, derive α_d from a pre-registered rule.
  4. [§4.3 and §3.8] There is an internal inconsistency in what 'Multi-hyp persona' means. Section 4.3 defines the condition as the 'best-scored hypothesis from Stage 5,' whereas the pipeline's stated contribution in Section 3.8 is confidence-weighted aggregation across M hypotheses via Eq. (5). The RQ1 results therefore do not evaluate the multi-hypothesis aggregation mechanism at all; they evaluate a single selected hypothesis. The authors should either evaluate the aggregate persona in RQ1 or explicitly restrict the multi-hypothesis claim to RQ3/RQ4.
  5. [§3.2, §3.6.3] The load-bearing premise that LLM-generated belief statements faithfully reflect the user's actual decision process is not tested. Section 3.6.3 says hallucination is 'constrained by construction,' but grounding a statement in a reconstructed percept does not guarantee that the statement matches what the user actually believed, and the paper cites known LLM Theory-of-Mind failures in Section 3.1. Without a fidelity check against the OPeRA interview rationales (which the dataset provides), the reconstructed percepts and personas could be coherent stereotypes. A human-annotation study on a sample of belief statements, or a comparison of inferred rationales against the interview rationales, would directly address this concern.
minor comments (5)
  1. [§4.2] The implementation details omit the embedding model's pooling and normalization choices for the MMR computation; adding these would improve reproducibility.
  2. [§4.4, Table 2] The Hit@3=100% result for the multi-hypothesis method on 27 users would benefit from a bootstrap confidence interval; with only 27 users, a single error changes the metric substantially.
  3. [§4.5, Table 4] The text says 'both methods achieve MAE=0.945' after listing Questionnaire and Direct Inference; consider stating the MAE once and explicitly noting that the two baselines coincide, to avoid the appearance of an error.
  4. [§5] The VisionOS case study is labeled a proof-of-concept, but the caption claims that 'the persona determines' the layout; softening this to 'is used to determine' or reporting a user study would prevent overstatement.
  5. [Throughout] Several hyperparameter values (λ=0.3, θ_dedup=0.85, K=50, M=7, temperature 0.85) are given without ablations or sensitivity checks; a short paragraph discussing their stability would strengthen the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No material circularity: external benchmarks anchor the claims; RQ1's format confound is a stated limitation, not a reduction.

full rationale

The derivation chain is not circular: each evaluation uses external ground truth rather than the pipeline's own outputs. RQ1 compares against OPeRA interview-based GT personas using a shared session prefix; RQ2 uses an explicit 70/30 session holdout with categories extracted from held-out interactions; RQ3 and RQ4 compare against self-reported Big Five and shopping-attitude surveys. Pipeline hyperparameters are fixed (M=7, lambda=0.3, theta_dedup=0.85, K=50), and the only prior is published Big Five population norms, so no fitted parameter is renamed as a prediction. The paper itself flags the main threat to RQ1 in Section 4.7, noting that the +50.6% margin 'may partly reflect prompt-format alignment' and that a format-controlled comparison is future work; this is a confound in the headline comparison, not a circular reduction. Eq. 5's alpha_d is under-specified, but the paper does not state that it was calibrated on the evaluation labels, so prior-shrinkage speculation would not meet the evidence bar. The only self-referential wording is Section 3.6.3's 'grounded in a reconstructed percept,' where the percept is itself an LLM reconstruction; this weakens the faithfulness claim but does not make any predicted number equivalent to an input. No load-bearing self-citation, imported uniqueness theorem, or ansatz-smuggled-via-citation appears.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests on treating LLM judgments as priors over human mental states, on OPeRA ground truth, and on population norms; no new physical or ontological entities are postulated. The 'persona' and 'beliefs' are interpretive constructs, not entities with independent falsifiable handles.

free parameters (6)
  • alpha_d (adaptive prior strength) = not disclosed
    Introduced in Eq. 5 to weight population priors against evidence; the mapping from evidence confidence to alpha_d is not specified, leaving room for hand-tuning that affects the stated Big Five MAE.
  • M (number of persona hypotheses) = 7
    Selected by hand; the ablation only compares M=1 vs M=7, so the reported gain may depend on this choice.
  • lambda (MMR diversity weight) = 0.3
    Hand-chosen in Eq. 4; affects which beliefs survive Stage 4.
  • theta_dedup (dedup cosine threshold) = 0.85
    Hand-chosen threshold in Stage 4 for collapsing near-duplicate beliefs.
  • K (max selected beliefs per user) = 50
    Hand-chosen cap on belief set size.
  • temperature (LLM decoding) = 0.85
    Decoding temperature for multi-hypothesis generation; influences diversity and quality.
assumptions (5)
  • domain assumption Rational-agent assumption in Eq. 1: observed actions are explainable by user beliefs given the perceived decision context.
    Invoked in Section 3.2; if false, backward belief inference has no formal basis.
  • domain assumption The LLM acts as an amortized approximate inference engine whose internalized priors over human beliefs are valid for this population.
    Stated in Section 3.2; the pipeline makes no likelihood computation and trusts LLM counterfactual reasoning.
  • domain assumption Reconstructed percepts from condensed HTML and metadata faithfully represent what the user saw, including visible alternatives.
    Stage 2, Section 3.5; viewport, scroll, and rendered state are not guaranteed in the trace.
  • domain assumption OPeRA ground truth, including personality assessments, attitude surveys, and interview personas, is valid for evaluation.
    Section 4.1; all four research questions measure against this data.
  • domain assumption Published Big Five population norms are appropriate priors for this 46-user cohort.
    Used in Eq. 5; if norms do not match the cohort, shrinkage targets are wrong.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces." pith.science (2026). https://pith.science/paper/W4FOM4OJ

@misc{pith2026260811354,
  author       = {Pith},
  title        = {Pith review of: Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W4FOM4OJ}},
  note         = {Machine review of arXiv:2608.11354}
}
read the original abstract

Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.

Figures

Figures reproduced from arXiv: 2608.11354 by the authors.

Figure 1
Figure 1. Big Five personality prediction results (N=46). Mean [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Per-item shopping attitude accuracy (N=46). Our [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Persona-driven spatial banking prototype on Vi [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 56 canonical work pages

  1. [1]

    Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B

    Chris L. Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B. Tenenbaum

  2. [2]

    Baker, Rebecca Saxe, and Joshua B

    Chris L. Baker, Rebecca Saxe, and Joshua B. Tenenbaum. 2009. Action Under- standing as Inverse Planning.Cognition113, 3 (2009), 329–349

  3. [3]

    Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation. InProceedings of the 17th ACM Conference on Recommender Systems. 1007–1014

  4. [4]

    Theory of Mind

    Simon Baron-Cohen, Alan M. Leslie, and Uta Frith. 1985. Does the Autistic Child Have a “Theory of Mind”?Cognition21, 1 (1985), 37–46

  5. [5]

    Pranav Bhandari, Usman Naseem, Amitava Datta, Nicolas Fay, and Mehwish Nasim. 2025. Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires. InCompanion Proceedings of the ACM on Web Conference 2025. ACM, 868–872

  6. [6]

    Marcel Binz and Eric Schulz. 2023. Turning large language models into cognitive models.arXiv preprint arXiv:2306.03917(2023)

  7. [7]

    Jaime Carbonell and Jade Goldstein. 1998. The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries. InProceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 335–336

  8. [8]

    Lalor, Yi Yang, and Ahmed Abbasi

    Shuai Chen, John P. Lalor, Yi Yang, and Ahmed Abbasi. 2025. PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins. InProceedings of the Fourth Workshop on Generation, Evaluation, and Metrics

Show all 95 references
  1. [9]

    Xiaocong Chen, Lina Yao, Aixin Sun, Xianzhi Wang, Xiwei Xu, and Liming Zhu

  2. [10]

    Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, et al. 2024. ToMBench: Benchmarking Theory of Mind in Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computat...

  3. [11]

    Konstantina Christakopoulou, Alberto Lalama, Cj Adams, Iris Qu, Tal Aber, Daiyi Azzato, Sasa Biber, Stephanie Buchner, Kellie Bui, Mauri Byon, et al. 2023. Large Language Models for User Interest Journeys.arXiv preprint arXiv:2305.15498 (2023)

  4. [12]

    2015.Click Models for Web Search

    Aleksandr Chuklin, Ilya Markov, and Maarten de Rijke. 2015.Click Models for Web Search. Morgan & Claypool Publishers

  5. [13]

    Viviane Clay, Peter König, and Sabine U. König. 2019. Eye Tracking in Virtual Reality.Journal of Eye Movement Research12, 1 (2019)

  6. [14]

    Logan Cross, Violet Xiang, Aditya Bhatia, Daniel L. K. Yamins, and Nick Haber

  7. [15]

    Joost C. F. de Winter, Tom Driessen, and Dimitra Dodou. 2024. The Use of Chat- GPT for Personality Research: Administering Questionnaires Using Generated Personas.Personality and Individual Differences225 (2024), 112669

  8. [16]

    Shanshan Feng, Lucas Vinh Tran, Gao Cong, Lisi Chen, Jing Li, and Fan Li. 2020. HME: A Hyperbolic Metric Embedding Approach for Next-POI Recommendation. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 202...

  9. [17]

    Chongming Gao, Shijun Li, Wenqiang Lei, Jiawei Chen, Biao Li, Peng Jiang, Xiangnan He, Jiaxin Mao, and Tat-Seng Chua. 2022. KuaiRec: A Fully-observed Dataset and Insights for Evaluating Recommender Systems. InProceedings of the 31st ACM International Conference on Information ...

  10. [18]

    Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5). InProceedings of the 16th ACM Conference on Recommender Systems. 299–315

  11. [19]

    Gershman and Noah D

    Samuel J. Gershman and Noah D. Goodman. 2014. Amortized inference in probabilistic reasoning.Proceedings of the Annual Meeting of the Cognitive Science Society36 (2014)

  12. [20]

    Matej Gjurković, Mladen Karan, Iva Vukojević, Mihaela Bošnjak, and Jan Šnajder

  13. [21]

    Jens Grubert, Tobias Langlotz, Stefanie Zollmann, and Holger Regenbrecht. 2017. Towards Pervasive Augmented Reality: Context-Awareness in Augmented Reality. IEEE Transactions on Visualization and Computer Graphics23, 6 (2017), 1706–1724

  14. [22]

    Huifeng Guo, Jinkai Yu, Qing Liu, Ruiming Tang, and Yuzhou Zhang. 2019. Buying or Browsing?: Predicting Real-time Purchasing Intent using Attention-based Deep Network with Multiple Behavior. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery a...

  15. [23]

    Wei Guo, Yaochen Lu, Yigang Li, Jiaqi Yang, Yingxiang Tan, Xingyu Zhang, Mingyang Chen, Junwei Zhang, Zhicheng Su, Weiwen Zhang, Yichao Chen, Bo Zheng, and Ruiming Tang. 2025. LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders.arXiv preprint arXiv:2504.12561(2025)

  16. [24]

    Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural Collaborative Filtering. InProceedings of the 26th International Conference on World Wide Web. 173–182

  17. [25]

    Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley. 2023. Large Language Models as Zero-Shot Conversational Recommenders. InProceedings of the 32nd ACM International Conference on Informatio...

  18. [26]

    Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk

  19. [27]

    Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2024. Large Language Models are Zero-Shot Rankers for Recommender Systems. InProceedings of the 46th European Conference on Information Retrieval. Springer, 364–381

  20. [28]

    Yifan Hu, Yehuda Koren, and Chris Volinsky. 2008. Collaborative Filtering for Implicit Feedback Datasets. InProceedings of the 8th IEEE International Conference on Data Mining. IEEE, 263–272

  21. [29]

    Minzhi Huang, Xiaolin Zhang, Christopher Soto, and James Evans. 2026. Design- ing AI-Agents with Personalities: A Psychometric Approach.Personality Science (2026)

  22. [30]

    Julian Jara-Ettinger. 2019. Theory of Mind as Inverse Reinforcement Learning. Current Opinion in Behavioral Sciences29 (2019), 105–110

  23. [31]

    Aditya Joshi, Asim Ahmad, and Ashutosh Modi. 2024. COLD: Causal reasOning in cLosed Daily Activities. InAdvances in Neural Information Processing Systems, Vol. 37

  24. [32]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recom- mendation. InProceedings of the IEEE International Conference on Data Mining. IEEE, 197–206

  25. [33]

    Junseok Kim, Nakyeong Yang, and Kyomin Jung. 2025. Persona is a Double-Edged Sword: Rethinking the Impact of Role-play Prompts in Zero-shot Reasoning Tasks. InFindings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-...

  26. [34]

    Yehuda Koren, Robert Bell, and Chris Volinsky. 2009. Matrix Factorization Tech- niques for Recommender Systems.Computer42, 8 (2009), 30–37

  27. [35]

    Michal Kosinski. 2024. Evaluating Large Language Models in Theory of Mind Tasks.Proceedings of the National Academy of Sciences121, 45 (2024), e2405460121

  28. [36]

    Chenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita Turakhia, Sonia Castelo Quispe, and Dong Li. 2025. Satori: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, 1–24. ...

  29. [37]

    Huao Li, Yu Chong, Simon Stepputtis, Joseph Campbell, Dana Hughes, Charles Lewis, and Subbarao Ramamoorthy. 2023. Theory of Mind for Multi-Agent Collaboration via Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 180–192

  30. [38]

    Jing Li, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma. 2017. Neural Attentive Session-based Recommendation. InProceedings of the 2017 ACM on Conference on Information and Knowledge Management. 1419–1428

  31. [39]

    Renhao Li, Heyang Xia, Xiang Yuan, Qingxiu Dong, Lei Sha, Zhixing Li, and Zhi- fang Liu. 2025. How Far Are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation. InFindings of the Association for Computational Linguistics: ACL 2025

  32. [40]

    Yuanchun Li, Hao Hao, Yizhi Ge, Haoyu Yu, and Yunxin Liu. 2024. Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.arXiv preprint arXiv:2401.05459(2024)

  33. [41]

    David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-Aware Online Adaptation of Mixed Reality Interfaces. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology. ACM, 147–160

  34. [42]

    Tian Liu, Yueming Xu, Jie Yang, and Kexin Li. 2025. Mindful Machines: Under- standing How AI’s Theory of Mind Capabilities Influence Consumer Response to Product Recommendations.Psychology & Marketing(2025). doi:10.1002/mar. 70022

  35. [43]

    Chris Lo, Dan Frankowski, and Jure Leskovec. 2016. Understanding the Rela- tionship between Browsing and Purchasing in a Large-Scale E-Commerce Site. InProceedings of the 25th ACM International on Conference on Information and Knowledge Management. 2421–2426

  36. [44]

    Feiyu Lu, Mengyu Chen, Hsiang Hsu, Pranav Deshpande, Cheng Yao Wang, and Blair MacIntyre. 2024. Adaptive 3D UI Placement in Mixed Reality Using Deep Reinforcement Learning. InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems. ACM, Article 32, 7 pag...

  37. [45]

    Hanjia Lyu, Song Jiang, Hanqing Zeng, Yinglong Xia, Qifan Wang, Si Zhang, Ren Chen, Christopher Leung, Jiajie Tang, and Jiebo Luo. 2023. LLM-Rec: Personal- ized Recommendation via Prompting Large Language Models.arXiv preprint arXiv:2307.15780(2023)

  38. [46]

    Yusen Mao, Shuangpeng Liu, Qiang Ni, Xudong Lin, and Liang He. 2024. A Review on Machine Theory of Mind.IEEE Transactions on Neural Networks and Inverse Theory of Mind Modeling for Content Recommendation Conference’17, July 2017, Washington, DC, USA Learning Systems(2024)

  39. [47]

    Elena Marchiori, Evangelos Niforatos, and Loris Preto. 2018. Analysis of Users’ Heart Rate Data and Self-Reported Perceptions to Understand Effective Virtual Reality Characteristics.Information Technology & Tourism18 (2018), 133–155

  40. [48]

    Sheshera Mysore, Zhuoran Lu, Mengting Wan, Tara Safavi, Jennifer Metzler, Julian McAuley, and Hamed Zamani. 2024. PEARL: Personalizing Large Language Model Writing Assistants via Extraction and Retrieval of User Profiles. InFindings of the Association for Computational Linguis...

  41. [49]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–22

  42. [50]

    Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, and Michael S

    Joon Sung Park, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, and Michael S. Bernstein. 2024. Generative Agent Simulations of 1,000 People.arXiv preprint arXiv:2411.10109(2024)

  43. [51]

    Shrivastva Viraj Pawar, Bharadwaj Pedapudi, and Pragya Kaushik. 2025. EARL: Early Intent Recognition in GUI Tasks Using Theory of Mind. InICML 2025 Workshop on Foundation Models in the Wild. https://openreview.net/forum?id= nABg9kZ7JR

  44. [52]

    Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier

    Max Pellert, Clemens M. Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. 2024. AI Psychometrics: Assessing the Psychological Profiles of Large Language Models Through Psychometric Inventories.Perspectives on Psychological Science19, 5 (2024), 808–826

  45. [53]

    Alexander Peysakhovich. 2019. Reinforcement Learning and Inverse Reinforce- ment Learning with System 1 and System 2. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 409–415

  46. [54]

    David Premack and Guy Woodruff. 1978. Does the Chimpanzee Have a Theory of Mind?Behavioral and Brain Sciences1, 4 (1978), 515–526

  47. [55]

    Rabinowitz, Frank Perbet, H

    Neil C. Rabinowitz, Frank Perbet, H. Francis Song, Chiyuan Zhang, S. M. Ali Eslami, and Matthew Botvinick. 2018. Machine Theory of Mind. InProceedings of the 35th International Conference on Machine Learning. 4218–4227

  48. [56]

    Filip Radlinski, Krisztian Balog, Bill Byrne, and Karthik Krishnamoorthi. 2022. On Natural Language User Profiles for Transparent and Scrutable Recommendation. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SI...

  49. [57]

    Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q

    Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan H. Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Saber, Manzil Zaheer, Ed H. Chi, and Jonathon Shlens. 2023. Recommender Systems with Generative Retrieval. InAdvances in Neural Information Pro...

  50. [58]

    Miquel Ramírez and Hector Geffner. 2010. Probabilistic Plan Recognition Us- ing Off-the-Shelf Classical Planners. InProceedings of the AAAI Conference on Artificial Intelligence. 1121–1126

  51. [59]

    Rahmani, Xi Wang, Xiao Fu, and Aldo Lipani

    Jerome Ramos, Hossein A. Rahmani, Xi Wang, Xiao Fu, and Aldo Lipani. 2024. Transparent and Scrutable Recommendations Using Natural Language User Profiles. InProceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (ACL ’24). ACL

  52. [60]

    Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme

  53. [61]

    Adam Richardson and Hari Srinivasan. 2024. User-LLM: Efficient LLM Contextu- alization with User Embeddings.arXiv preprint arXiv:2402.13598(2024)

  54. [62]

    Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2022. Neural Theory- of-Mind? On the Limits of Social Intelligence in Large LMs. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 3762–3780

  55. [63]

    Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. 2024. Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models. InPro- ceedings of the 18th Conference of the European ...

  56. [64]

    Herbert A. Simon. 1955. A Behavioral Model of Rational Choice.The Quarterly Journal of Economics69, 1 (1955), 99–118

  57. [65]

    James W. A. Strachan, Dalila Parga Pelaéz, Gyles W. Humphreys, Michel Thiebaut de Schotten, and Giosuè Baggio. 2024. Testing Theory of Mind in Large Language Models and Humans.Nature Human Behaviour8 (2024), 1285– 1295

  58. [66]

    Winnie Street, John Oliver Siy, Geoff Keeling, Adrien Baranes, Benjamin Barnett, Michael McKibben, Tatenda Kanyere, Alison Lentz, Blaise Aguera y Arcas, and Robin Moran. 2024. LLMs Achieve Adult Human Performance on Higher-Order Theory of Mind Tasks.arXiv preprint arXiv:2405.1...

  59. [67]

    Andreas Stuhlmüller, Jacob Taylor, and Noah D. Goodman. 2013. Learning stochastic inverses. InAdvances in Neural Information Processing Systems, Vol. 26

  60. [68]

    Gita Sukthankar, Christopher Geib, Hung Hai Bui, David Pynadath, and Robert P. Goldman. 2014.Plan, Activity, and Intent Recognition: Theory and Practice. Morgan Kaufmann

  61. [69]

    Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang

  62. [70]

    Amanda Swearngin, Mira Dontcheva, Wilmot Li, Joel Brandt, Morgan Dixon, and Andrew J. Ko. 2018. Rewire: Interface Design Assistance from Examples. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM, 1–12

  63. [71]

    Lucas Vinh Tran, Yi Tay, Shuai Zhang, Gao Cong, and Xiaoli Li. 2020. HyperML: A Boosting Metric Learning Approach in Hyperbolic Space for Recommender Systems. InWSDM ’20: The Thirteenth ACM International Conference on Web Search and Data Mining, Houston, TX, USA, February 3-7,...

  64. [72]

    Tomer Ullman. 2023. Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.arXiv preprint arXiv:2302.08399(2023)

  65. [73]

    Simine Vazire and Matthias R. Mehl. 2008. Knowing Me, Knowing You: The Accuracy and Unique Predictive Validity of Self-Ratings and Other-Ratings of Daily Behavior.Journal of Personality and Social Psychology95, 5 (2008), 1202– 1216

  66. [74]

    Ones, Lihong He, and Xiaolin Xu

    Yuting Wang, Jiayi Zhao, Deniz S. Ones, Lihong He, and Xiaolin Xu. 2025. Evalu- ating the Ability of Large Language Models to Emulate Personality.Scientific Reports15 (2025), 1–15

  67. [75]

    Zhenlei Wang, Xu Chen, Rui Zhou, Quanyu Dai, Zhenhua Dong, and Ji-Rong Wen. 2023. Sequential Recommendation with Causal Behavior Discovery. In Proceedings of the IEEE International Conference on Data Engineering (ICDE). 1849–1862

  68. [76]

    Ziyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, and Dakuo Wang. 2025. OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evalua...

  69. [77]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.Advances in Neural Information Processing Systems35 (2022), 24824–24837

  70. [78]

    Henry M. Wellman. 2014.Making Minds: How Theory of Mind Develops. Oxford University Press

  71. [79]

    Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. 2020. MIND: A Large-scale Dataset for News Recommendation. InProceedings of the 58th Annual Meeting of the Association for Computational Ling...

  72. [80]

    Pynadath

    Haochen Wu, Pedro Sequeira, and David V. Pynadath. 2023. Multiagent In- verse Reinforcement Learning via Theory of Mind Reasoning.arXiv preprint arXiv:2302.10238(2023)

  73. [81]

    Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, Hui Xiong, and Enhong Chen. 2024. A Survey on Large Language Models for Recommendation.World Wide Web27, 5 (2024), 60

  74. [82]

    Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan. 2019. Session-based Recommendation with Graph Neural Networks. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 346–353

  75. [83]

    An Zhang, Yuxin Chen, Leheng Sheng, Xiang Wang, and Tat-Seng Chua. 2024. On Generative Agents in Recommendation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1807–1817. doi:10.1145/3626772.3657844

  76. [84]

    Junjie Zhang, Ruobing Xie, Yupeng Hou, Wayne Xin Zhao, Leyu Lin, and Ji-Rong Wen. 2023. Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach.arXiv preprint arXiv:2305.07001(2023)

  77. [85]

    Yimeng Zhang, Jiri Gesi, Ran Xue, Tian Wang, Ziyi Wang, Yuxuan Lu, Sinong Zhan, Huimin Zeng, Qingjun Cui, Yufan Guo, Jing Huang, Mubarak Shah, and Dakuo Wang. 2025. See, Think, Act: Online Shopper Behavior Simulation with VLM Agents.arXiv preprint arXiv:2510.19245(2025)

  78. [86]

    Mengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li, Pengjie Gu, Zhenhua Dong, Ruiming Tang, and Yi Cai. 2025. MTRec: Learning to Align with User Pref- erences via Mental Reward Models. InAdvances in Neural Information Processing Systems

  79. [87]

    Bowen Zheng, Yupeng Hou, Hongyu Lu, Zhichao Chen, Wayne Xin Zhao, and Ji- Rong Wen. 2024. Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation. InProceedings of the IEEE 40th International Conference on Data Engineering. IEEE, 1–14

  80. [88]

    Yaochen Zhu, Liang Wu, Qi Guo, Liangjie Hong, and Jundong Li. 2024. Collabo- rative Large Language Model for Recommender Systems. InProceedings of the ACM Web Conference 2024. 3162–3172

  81. [2009]

    InProceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence

    BPR: Bayesian Personalized Ranking from Implicit Feedback. InProceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence. 452–461

  82. [2016]

    In Proceedings of the 4th International Conference on Learning Representations

    Session-based Recommendations with Recurrent Neural Networks. In Proceedings of the 4th International Conference on Learning Representations

  83. [2017]

    Conference’17, July 2017, Washington, DC, USA Chen et al

    Rational Quantitative Attribution of Beliefs, Desires and Percepts in Human Mentalizing.Nature Human Behaviour1 (2017), 0064. Conference’17, July 2017, Washington, DC, USA Chen et al

  84. [2019]

    InProceedings of the 28th ACM International Conference on Information and Knowledge Management

    BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformers. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. 1441–1450

  85. [2020]

    Generative Inverse Deep Reinforcement Learning for Online Recommen- dation.arXiv preprint arXiv:2011.02248(2020)

  86. [2021]

    InProceedings of the Ninth International Workshop on Natural Language Processing for Social Media

    PANDORA Talks: Personality and Demographics on Reddit. InProceedings of the Ninth International Workshop on Natural Language Processing for Social Media. 138–152

  87. [2024]

    Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models.arXiv preprint arXiv:2407.07086(2024)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.