Pith. sign in

REVIEW 5 cited by

Can Large Language Model Agents Simulate Human Trust Behavior?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04559 v4 pith:FSIRBHDH submitted 2024-02-07 cs.AI cs.CLcs.HC

classification cs.AIcs.CLcs.HC
keywords trustagentsbehaviorhumanagenthumansmodelsimulate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Model (LLM) agents have been increasingly adopted as simulation tools to model humans in social science and role-playing applications. However, one fundamental question remains: can LLM agents really simulate human behavior? In this paper, we focus on one critical and elemental behavior in human interactions, trust, and investigate whether LLM agents can simulate human trust behavior. We first find that LLM agents generally exhibit trust behavior, referred to as agent trust, under the framework of Trust Games, which are widely recognized in behavioral economics. Then, we discover that GPT-4 agents manifest high behavioral alignment with humans in terms of trust behavior, indicating the feasibility of simulating human trust behavior with LLM agents. In addition, we probe the biases of agent trust and differences in agent trust towards other LLM agents and humans. We also explore the intrinsic properties of agent trust under conditions including external manipulations and advanced reasoning strategies. Our study provides new insights into the behaviors of LLM agents and the fundamental analogy between LLMs and humans beyond value alignment. We further illustrate broader implications of our discoveries for applications where trust is paramount.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research

    cs.MA 2025-08 conditional novelty 6.0 of 10

    Six LLMs show an equivalence-versus-process paradox: some match human surface behaviors but few replicate human decision pathways, so GABMs need dual-level validation before use in logistics research.

  2. SocialEval: Evaluating Social Intelligence of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SocialEval is a 153-tree bilingual benchmark that evaluates LLM social intelligence through goal outcomes and interpersonal ability choices, finding LLMs below humans and biased toward prosocial behavior.

  3. Exploration on Demand: From Algorithmic Control to User Empowerment

    cs.IR 2025-07 conditional novelty 4.0 of 10

    A user-controlled exploration layer over semantic movie clusters reduces recommendation redundancy (ILS 0.34 to 0.26) but collapses relevance (NDCG 0.00), earning preference from simulated long-history LLM users in 72...

  4. AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need

    cs.CL 2025-06 reject novelty 4.0 of 10

    A divide-and-conquer multi-agent framework with task forests and specialized roles improves math and code benchmarks but not commonsense or domain QA, and the adaptive heterogeneous-LLM engine is never tested.

  5. A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models

    cs.HC 2025-06 unverdicted novelty 2.0 of 10

    A review of human error research concluding that integrating AI and cognitive models into human reliability assessment can markedly improve predictive fidelity, but data scarcity and opacity remain barriers.

Pith tools