REVIEW 5 cited by
Explaining Large Language Models Decisions Using Shapley Values
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The emergence of large language models (LLMs) has opened up exciting possibilities for simulating human behavior and cognitive processes, with potential applications in various domains, including marketing research and consumer behavior analysis. However, the validity of utilizing LLMs as stand-ins for human subjects remains uncertain due to glaring divergences that suggest fundamentally different underlying processes at play and the sensitivity of LLM responses to prompt variations. This paper presents a novel approach based on Shapley values from cooperative game theory to interpret LLM behavior and quantify the relative contribution of each prompt component to the model's output. Through two applications - a discrete choice experiment and an investigation of cognitive biases - we demonstrate how the Shapley value method can uncover what we term "token noise" effects, a phenomenon where LLM decisions are disproportionately influenced by tokens providing minimal informative content. This phenomenon raises concerns about the robustness and generalizability of insights obtained from LLMs in the context of human behavior simulation. Our model-agnostic approach extends its utility to proprietary LLMs, providing a valuable tool for practitioners and researchers to strategically optimize prompts and mitigate apparent cognitive biases. Our findings underscore the need for a more nuanced understanding of the factors driving LLM responses before relying on them as substitutes for human subjects in survey settings. We emphasize the importance of researchers reporting results conditioned on specific prompt templates and exercising caution when drawing parallels between human behavior and LLMs.
Forward citations
Cited by 5 Pith papers
-
LLM-Based Social Simulations Require a Boundary
LLM-based social simulations are scientifically useful only within boundaries set by behavioral variance, and current validation practice under-checks variance.
-
Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions
GPT-3.5 and GPT-4 approximate human immigration preferences in a discrete choice experiment but exhibit systematic biases toward privileged nationalities and occupations.
-
SCAR: Shapley Credit Assignment for More Efficient RLHF
SCAR redistributes the terminal RLHF reward to tokens and spans via Shapley values, preserving the total return while improving training efficiency and final reward across three LLM alignment tasks.
-
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.
-
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
Frontier LLMs exhibit framing, anchoring, round-number, representativeness, and priming biases, and word-level Shapley attribution can localize the words that drive those biases.
Discussion (0). Continue with ORCID to comment.