Pith. sign in

REVIEW 5 cited by

Explaining Large Language Models Decisions Using Shapley Values

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01332 v3 pith:JS4LJDQ6 submitted 2024-03-29 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords behaviorhumanllmscognitivepromptshapleyapplicationsapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The emergence of large language models (LLMs) has opened up exciting possibilities for simulating human behavior and cognitive processes, with potential applications in various domains, including marketing research and consumer behavior analysis. However, the validity of utilizing LLMs as stand-ins for human subjects remains uncertain due to glaring divergences that suggest fundamentally different underlying processes at play and the sensitivity of LLM responses to prompt variations. This paper presents a novel approach based on Shapley values from cooperative game theory to interpret LLM behavior and quantify the relative contribution of each prompt component to the model's output. Through two applications - a discrete choice experiment and an investigation of cognitive biases - we demonstrate how the Shapley value method can uncover what we term "token noise" effects, a phenomenon where LLM decisions are disproportionately influenced by tokens providing minimal informative content. This phenomenon raises concerns about the robustness and generalizability of insights obtained from LLMs in the context of human behavior simulation. Our model-agnostic approach extends its utility to proprietary LLMs, providing a valuable tool for practitioners and researchers to strategically optimize prompts and mitigate apparent cognitive biases. Our findings underscore the need for a more nuanced understanding of the factors driving LLM responses before relying on them as substitutes for human subjects in survey settings. We emphasize the importance of researchers reporting results conditioned on specific prompt templates and exercising caution when drawing parallels between human behavior and LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Based Social Simulations Require a Boundary

    cs.CY 2025-06 conditional novelty 6.0 of 10

    LLM-based social simulations are scientifically useful only within boundaries set by behavioral variance, and current validation practice under-checks variance.

  2. Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-3.5 and GPT-4 approximate human immigration preferences in a discrete choice experiment but exhibit systematic biases toward privileged nationalities and occupations.

  3. SCAR: Shapley Credit Assignment for More Efficient RLHF

    cs.AI 2025-05 conditional novelty 5.0 of 10

    SCAR redistributes the terminal RLHF reward to tokens and spans via Shapley values, preserving the total return while improving training efficiency and final reward across three LLM alignment tasks.

  4. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

  5. CBEval: A framework for evaluating and interpreting cognitive biases in LLMs

    cs.CL 2024-12 reject novelty 4.0 of 10

    Frontier LLMs exhibit framing, anchoring, round-number, representativeness, and priming biases, and word-level Shapley attribution can localize the words that drive those biases.

Pith tools