Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A modular LLM agent architecture that supplies personal characteristics, priors, and external information can simulate household and expert macroeconomic expectations closely enough to match human survey distributions and human-like…

desk verdict Solid framework paper, but the missing prior-only baseline leaves the architecture's added value unproven; worth refereeing with that condition. read the letter →

arxiv 2505.17648 v5 pith:W62VVCI2 submitted 2025-05-23 econ.GN cs.AIq-fin.EC

classification econ.GNcs.AIq-fin.EC
keywords macroeconomicexpectationsLLMagentssurveyexperimentsexpectationformationlargelanguagemodelssimulationselectiverecallmental
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes a framework, called UNITE, for building LLM-based economic agents that simulate how households and experts form macroeconomic expectations in survey experiments. It claims that agents equipped with modules supplying personal characteristics, prior expectations, and external or professional information generate expectation distributions that closely match human survey data across three representative experiments, while also capturing qualitative patterns in open-ended reasoning that prompt-only foundation models miss. The authors report that simulated distributions are more homogeneous than human ones and therefore position the method as a complement to, not a replacement for, traditional surveys. If the claim holds, economists gain a low-cost, scalable way to pre-test survey designs and pre-estimate future expectation distributions.

What carries the argument

The central object is the LLM Agent built inside a five-step pipeline the paper calls UNITE: Construction, Initialization, Simulation, Pre-estimation, and Evaluation. Each agent combines a general-purpose large language model, treated as the 'brain,' with functional modules: a Personal Characteristics Module (PCM), a Prior Expectations & Perceptions Module (PEPM), a Social Media Information Module (SMIM), and, for experts, a Professional Background Module (PBM) and a Knowledge Acquisition Module (KAM). Initialization prompts assign each agent a role, a confidence level, a task, and a rule for trading off priors against external signals, while random normal draws on the temperature and top-p sampling parameters stand in for unobserved human heterogeneity. The framework evaluates success by comparing histogram-based probability vectors of simulated and human expectations using Pearson correlation and cosine similarity, and by comparing the causal Directed Acyclic Graphs, or mental models, extracted from open-ended responses.

What would settle it

Run the same three experiments with agents that receive only the prior-expectation inputs, or a simple statistical transform of those priors, and compare the distributional similarity to the full-agent results; if the prior-only baseline matches human distributions as well as the full agents do, the claim that the architecture adds simulation power would be falsified.

Watch

Extended reading notes

Core claim

The central discovery is that the architecture of the agent, rather than the underlying large language model alone, is what makes expectation simulation work. Across three benchmark designs—hypothetical vignette experiments in which households and experts react to stylized macroeconomic shocks, information-provision experiments on home price expectations, and an out-of-sample pre-estimation of the 2025 Michigan Survey of Consumers—the authors report that the shape similarity between simulated and human expectation distributions, measured by Pearson correlation and cosine similarity of histogram-based probability vectors, averages around 0.8 and rarely falls below 0.5. Ablations show that removing the Prior Expectations & Perceptions Module sharply degrades the distributional match, while removing personal-characteristics or social-media modules degrades the human-likeness of the thoughts and selective-recall patterns. Foundation models given only initialization prompts produce far less human-aligned distributions and mental models. The paper concludes that priors are the main driver of distributional fidelity, whereas personal, professional, and external-information modules are what establish human-like reasoning, and that the agents therefore narrow the belief gap between generative AI and humans at the aggregate level.

Load-bearing premise

The central claim rests on the assumption that the prior beliefs fed into the agents are not themselves almost enough to reproduce the human distributions, so the modules and initialization do real work beyond passing those priors through.

Editorial extensions

If this is right

  • Economists could use the LLM Agents to pre-test survey questionnaires and vignettes before fielding costly human surveys.
  • The same architecture can pre-estimate future expectation distributions, as demonstrated by the out-of-sample 2025 exercise, without waiting for survey data to be released.
  • The ablation results give design guidance: invest in prior-expectation data for distributional fidelity and in personal, social, and professional modules for human-like reasoning.
  • Because simulated distributions are more homogeneous than human ones, the framework should be treated as a complement or pilot, not a substitute for human subjects.
  • Prompt-only foundation models are not enough; the modular architecture and initialization are what carry the simulation performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The most direct test the authors leave implicit is a prior-only baseline: until agents fed only the prior-expectation inputs are compared with the full architecture, some of the credit for the distributional match could belong to the survey inputs rather than to the modules.
  • A natural extension would be to calibrate the random disturbance parameters to the observed dispersion of human responses, rather than using fixed normal draws, which could reduce the documented over-homogeneity.
  • The same modular construction could be carried over to other belief objects, such as stock market expectations, policy narratives, or firm price-setting, wherever priors and external information can be sourced.
  • The mental-model comparisons imply a testable connection: agents that produce more diverse causal graphs should also produce less homogeneous expectation distributions, linking the reasoning dimension to the distributional dimension.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper proposes a modular framework ('UNITE') for simulating macroeconomic expectations in survey experiments using LLM-based agents. The authors construct Household Agents equipped with a Personal Characteristics Module, a Prior Expectations & Perceptions Module, and a Social Media Information Module, and Expert Agents with a Professional Background Module, a Prior Expectations & Perceptions Module, and a Knowledge Acquisition Module. They validate the framework by replicating three survey designs: the Andre et al. (2022) hypothetical vignette experiments on inflation and unemployment expectations of households and experts, the Chopra et al. (2025) information-provision experiment on home price expectations of homeowners and renters, and a pre-estimation exercise for the 2025 Michigan Survey of Consumers. The evaluation uses distributional similarity metrics (Pearson correlation and cosine similarity) between simulated and human distributions, qualitative analysis of open-ended responses, and mental-model comparisons via causal DAGs. The paper concludes that LLM Agents generate expectation distributions highly similar to human data, outperform foundation models relying simply on prompt engineering, and that the prior expectations module is crucial for distributional matching while other modules drive human-like thought processes.

Significance. If the central claim were established, this framework would be a useful, low-cost, and scalable complement to traditional survey experiments on expectation formation. The paper has several strengths: it covers three distinct and representative experimental designs, includes a detailed modular architecture, provides an ablation study (Figure 12 and Appendix Figures A.20-A.23), and attempts an out-of-sample pre-estimation exercise for the 2025 MSC. The mental-model analysis via DAGs and the comparison with foundation models are also valuable contributions. However, the central distributional claim is currently under-identified. The Prior Expectations & Perceptions Module feeds each agent the contemporaneous priors of the very populations used as benchmarks, yet no baseline is reported that outputs or statistically transforms these priors directly into an expectation distribution. The ablation shows that removing PEPM degrades performance, but this only demonstrates that priors matter; it does not demonstrate that the LLM agent architecture adds value beyond those priors. This missing baseline is load-bearing for the paper's headline conclusion.

major comments (3)
  1. [Sections 3.1-3.2, 5.1 (Figure 5), 6 (Figure 12)] The validation of the distributional claim lacks a prior-only baseline. The PEPM supplies agents with contemporaneous prior expectations from the same populations that serve as benchmarks: 2019 MSC and SPF priors for the Andre et al. vignettes, 2024 Chopra et al. priors for the information-provision experiment, and 2024 MSC priors for the 2025 MSC pre-estimation. The ablation in Figure 12 shows that removing PEPM sharply reduces distributional similarity, but this establishes only that the injected priors are important, not that the modular LLM architecture contributes beyond those priors. A baseline that directly outputs (or applies a simple statistical transformation to) the prior distribution could plausibly match Figure 5 as well as the LLM agents do, particularly for the information-provision experiment where the posterior may be close to the prior. The authors should add such baselines for all three experiments and report the incremental improvement of LLM Agents over them. Without this, the claim that the framework 'simulates' expectations rather than merely re-issuing survey priors is not supported.
  2. [Section 4.3, Figure 5(d)] The out-of-sample MSC test covers a single period (2025) and uses 2024 MSC priors from the same survey. Since the 2024 prior distribution may be a strong forecast of the 2025 distribution, this design does not provide independent evidence for the architecture unless compared against a prior-only baseline. The authors should report, for example, the similarity between the 2024 MSC distribution (or a simple lagged/conditional version of it) and the 2025 MSC distribution, and show that the LLM agents' pre-estimation improves on that benchmark. A multi-period or multi-horizon out-of-sample evaluation would also strengthen the claim of pre-estimation capability.
  3. [Section 5.1, Figure 5] The distributional similarity metrics—Pearson correlation and cosine similarity on histogram-based probability vectors—are weak criteria for 'highly similar' distributions. High correlation can coexist with substantial mean shifts, and the values depend on the chosen binning rule (Freedman-Diaconis or approximate sample-size bins). The authors should supplement these metrics with more stringent comparisons such as the Kolmogorov-Smirnov statistic, Wasserstein distance, or direct comparisons of means and variances between simulated and human distributions. This is important because the headline claim of distributional fidelity rests entirely on these two similarity measures.
minor comments (4)
  1. [Section 7] In the concluding remarks, 'Extending the framework to stimulate the expectations of firms' should be 'simulate'; this typo also appears in the supplementary appendix headings.
  2. [Figure 5 caption] The description of the bootstrap confidence intervals ('obtained by bootstrap over histogram-based probability vectors') is unclear; please specify whether the bootstrap resamples respondents, agents, or histogram bins, and how the confidence intervals are constructed.
  3. [Supplementary Appendix Table A.1] The knowledge cutoffs for Qwen3-235B-A22B-Thinking-2507 and DeepSeek-R1-0528 are inferred by querying the models rather than officially disclosed; this is a reasonable approximation but should be presented with appropriate caution, as the contamination-freeness argument depends on these cutoffs.
  4. [Section 3.1, Random Disturbances] The random disturbance distributions for Temperature and Top-p are introduced as free parameters, but no sensitivity analysis is reported for alternative parameterizations; a brief robustness check or discussion of how sensitive the distributional similarity results are to these choices would be useful.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the simulated distributions are not identical to the PEPM inputs by construction, and the out-of-sample MSC exercise provides independent support.

full rationale

The paper's derivation chain does not reduce to its inputs by construction. In the hypothetical vignette experiments, the PEPM feeds categorical prior expectations from the 2019 MSC/SPF, but the benchmarked outcome is the shock-induced change in expectations, computed as the difference between shock-scenario and baseline-scenario forecasts. That target is a distinct quantity, not the prior itself. In the information-provision experiments, the PEPM inputs individual prior home-price expectations, while the target is a posterior distribution elicited after a 6% or 2% forecast treatment; the posterior is again a different object from the prior. The ablation in Figure 12 shows that removing PEPM degrades distributional similarity, which is an empirical component-contribution result rather than a definitional identity. The MSC pre-estimation exercise is genuinely out-of-sample: 2024 priors are used to simulate 2025 expectations for a survey released after the models' knowledge cutoffs, providing independent support. The paper does not rely on self-citations, author-imported uniqueness theorems, or ansatz smuggled in via citation. A missing prior-only baseline is a legitimate validity concern about whether the priors alone could reproduce the target distributions, but the hard circularity rules require exhibiting that a 'prediction' equals its input by construction; that requirement is not met here. Accordingly, the appropriate finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The framework relies on several researcher-chosen distributions (temperature, top-p), a predefined confidence rule, and domain assumptions about what drives expectations. The paper does not compare against a prior-only baseline, so the causal contribution of each module is inferred from ablations rather than a formal identification strategy.

free parameters (3)
  • Temperature distribution for random disturbances = N(1.0, 0.5)
    Chosen by hand to reflect unobserved heterogeneity; no fitting procedure described; affects diversity of simulated responses.
  • Top-p distribution for random disturbances = N(0.5, 0.25)
    Chosen by hand alongside temperature; controls candidate token set; no evidence that choices are not tuned to improve distributional match.
  • Five-level confidence assignment = random stratified assignment
    Household agents are randomly assigned one of five confidence levels; this changes how much they rely on priors vs. external information and is an ad hoc design choice.
assumptions (5)
  • domain assumption Household expectations are primarily shaped by personal characteristics, prior expectations, and social media information.
    Section 3.1 cites literature to justify PCM, PEPM, and SMIM; this is a modeling assumption that may not capture all relevant drivers.
  • domain assumption Expert expectations are primarily shaped by professional background and domain knowledge, with demographic characteristics playing a minor role.
    Section 3.2 states this based on Benchimol et al. (2022); the validity is assumed for the expert agent design.
  • ad hoc to paper The confidence-based trade-off rule between priors and signals accurately describes human belief updating.
    The initialization prompts instruct agents to overweight priors when confident and external information when not; this rule is not derived from data and is a key behavioral assumption.
  • ad hoc to paper LLM-generated synthetic expert profiles are similar enough to real profiles to serve as valid simulation inputs.
    The PBM uses synthetic data when real expert samples are small; the paper does not validate that these synthetic profiles produce unbiased expectations.
  • domain assumption The chosen foundation model, Qwen3-235B-A22B-Thinking-2507, can faithfully role-play households and experts given the provided context.
    The success of the framework depends on the LLM's instruction-following and role-playing abilities; this is an empirical premise.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents." pith.science (2026). https://pith.science/paper/W62VVCI2

@misc{pith2026250517648,
  author       = {Pith},
  title        = {Pith review of: Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W62VVCI2}},
  note         = {Machine review of arXiv:2505.17648}
}
read the original abstract

We introduce a framework for simulating macroeconomic expectations in survey experiments using LLM-based economic agents (LLM Agents). We construct LLM Agents equipped with several functional modules that retrieve personal characteristics, prior expectations, and dynamic external information. We validate our framework by recapitulating three representative survey designs covering various expectations across different types of respondents. Our results show that LLM Agents generate expectation distributions highly similar to human data and capture human-aligned qualitative patterns in open-ended responses. Evaluation reveals that priors are crucial for matching distributions, whereas personal and external information drive human-like thought processes. Our findings offer guidance for narrowing the belief gap between generative AI and humans at the aggregate level while delineating the boundaries of the framework.

Figures

Figures reproduced from arXiv: 2505.17648 by the authors.

Figure 1
Figure 1. The “UNITE” framework for simulating macroeconomic expectations [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. LLM Agents for simulating household expectations [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. LLM Agents for simulating expert expectations [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Overview of the experimental procedure and structure of th [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Shape similarity between the expectation distributions generated [PITH_FULL_IMAGE:figures/full_fig_p025_5.png]
Figure 6
Figure 6. Figure 6: Comparison of LLM Agents’ simulated results with human data in Sub-Exp [PITH_FULL_IMAGE:figures/full_fig_p025_6.png]
Figure 7
Figure 7. Figure 7: Word usage for open-ended responses of humans and LLM Agents acros [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: Response types in open-ended responses of humans and LLM Agent [PITH_FULL_IMAGE:figures/full_fig_p030_8.png]
Figure 9
Figure 9. Figure 9: Open-ended responses on how higher expected home price growth affe [PITH_FULL_IMAGE:figures/full_fig_p032_9.png]
Figure 10
Figure 10. Figure 10: Proportion of various channels recalled by Household Agents when [PITH_FULL_IMAGE:figures/full_fig_p033_10.png]
Figure 11
Figure 11. Figure 11: The “average” DAGs underlying the formation of unemployme [PITH_FULL_IMAGE:figures/full_fig_p036_11.png]
Figure 12
Figure 12. Figure 12: Shape similarity between the expectation [PITH_FULL_IMAGE:figures/full_fig_p038_12.png]
Figure 13
Figure 13. Figure 13: Response types in open-ended responses of LLM Agents ( [PITH_FULL_IMAGE:figures/full_fig_p040_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncovering Salience-Driven Dynamics in Consumer Confidence with Generative Social Simulation

    cs.CY 2026-06 unverdicted novelty 6.0 of 10

    ConsumerSim reconstructs official CCI series from synthetic populations and multi-source signals, outperforming baselines on reconstruction metrics and aiding short-horizon activity predictions.

  2. OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

    cs.AI 2026-03 conditional novelty 5.0 of 10

    OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.

Reference graph

Works this paper leans on

104 extracted references · 69 canonical work pages · cited by 2 Pith papers

  1. [1]

    The Review of Economic Studies , year =

    Narratives about the Macroeconomy , author =. The Review of Economic Studies , year =. doi:10.1093/restud/rdag014 , note =

  2. [2]

    and Pizzinelli, C

    Andre, P. and Pizzinelli, C. and Roth, C. and Wohlfart, J. , year =. Subjective Models of the Macroeconomy: Evidence From Experts and Representative Samples , journal =

  3. [3]

    and Schirmer, P

    Andre, P. and Schirmer, P. and Wohlfart, J. , year =. Mental Models of the Stock Market , journal =

  4. [4]

    and Marcucci, J

    Angelico, C. and Marcucci, J. and Miccoli, M. and Quarta, F. , year =. Can we measure inflation expectations using Twitter? , journal =

  5. [5]

    Argyle, L. P. and Busby, E. C. and Fulda, N. and Gubler, J. R. and Rytting, C. and Wingate, D. , year =. Out of One, Many: Using Language Models to Simulate Human Samples , journal =

  6. [6]

    and El-Shagi, M

    Benchimol, J. and El-Shagi, M. and Saadon, Y. , year =. Do expert experience and characteristics affect inflation forecasts? , journal =

  7. [7]

    and Fermand, E

    Ben-David, I. and Fermand, E. and Kuhnen, C. and Li, G. , year =. Expectations Uncertainty and Household Economic Behavior , institution =

  8. [8]

    Benjamin, D. J. , year =. Errors in probabilistic reasoning and judgment biases , booktitle =

Show all 104 references
  1. [9]

    and Kuang, P

    Binder, C. and Kuang, P. and Tang, L. , year =. Central Bank Communication and House Price Expectations , institution =

  2. [10]

    Journal of the European Economic Association , volume =

    Binder, Carola and Kuang, Pei and Tang, Li , title =. Journal of the European Economic Association , volume =. 2026 , month =

  3. [11]

    and Cong, L

    Bini, P. and Cong, L. and Huang, X. and Jin, L. J. , year =. Behavioral Economics of. SSRN Electronic Journal , doi =

  4. [12]

    and Burro, G

    Bordalo, P. and Burro, G. and Coffman, K. and Gennaioli, N. and Shleifer, A. , year =. Imagining the Future: Memory, Simulation, and Beliefs , journal =

  5. [13]

    and Coffman, K

    Bordalo, P. and Coffman, K. and Gennaioli, N. and Shleifer, A. , year =. Stereotypes , journal =

  6. [14]

    and Eriksen, J

    Bro, J. and Eriksen, J. N. , year =. Subjective expectations and house prices , journal =

  7. [15]

    and D'Acunto, F

    Bruschi, C. and D'Acunto, F. and Kumar, S. and Weber, M. , year =. Subjective Models of Workers and Managers for Macroeconomic Expectations , journal =

  8. [16]

    Bybee, J. L. , year =. The Ghost in the Machine: Generating Beliefs with Large Language Models , institution =

  9. [17]

    Carroll, C. D. , year =. Macroeconomic Expectations of Households and Professional Forecasters , journal =

  10. [18]

    and Charness, G

    Chan, K. and Charness, G. and Dave, C. and Reddinger, J. L. , year =. On Prior Confidence and Belief Updating , doi =

  11. [19]

    and Fang, H

    Chen, Y. and Fang, H. and Zhao, Y. and Zhao, Z. , year =. Recovering Overlooked Information in Categorical Variables with LLMs: An Application to Labor Market Mismatch , institution =

  12. [20]

    and Liu, T

    Chen, Y. and Liu, T. X. and Shan, Y. and Zhong, S. , year =. The emergence of economic rationality of GPT , journal =

  13. [21]

    and Roth, C

    Chopra, F. and Roth, C. and Wohlfart, J. , year =. Home Price Expectations and Spending: Evidence from a Field Experiment , journal =

  14. [22]

    and Gorodnichenko, Y

    Coibion, O. and Gorodnichenko, Y. , year =. Information Rigidity and the Expectations Formation Process: A Simple Framework and New Facts , journal =

  15. [23]

    and Gorodnichenko, Y

    Coibion, O. and Gorodnichenko, Y. and Kamdar, R. , year =. The Formation of Expectations, Inflation, and the Phillips Curve , journal =

  16. [24]

    and Gorodnichenko, Y

    Coibion, O. and Gorodnichenko, Y. and Kumar, S. and Pedemonte, M. , year =. Inflation expectations as a policy tool? , journal =

  17. [25]

    and Gorodnichenko, Y

    Coibion, O. and Gorodnichenko, Y. and Weber, M. , year =. Monetary Policy Communications and Their Effects on Household Inflation Expectations , journal =

  18. [26]

    and Li, N

    Cui, Z. and Li, N. and Zhou, H. , year =. A large-scale replication of scenario-based experiments in psychology and management using large language models , journal =

  19. [27]

    Curtin, R. T. , year =. Indicators of Consumer Behavior: The University of Michigan Surveys of Consumers , journal =

  20. [28]

    and Charalambakis, E

    D'Acunto, F. and Charalambakis, E. and Georgarakos, D. and Kenny, G. and Meyer, J. and Weber, M. , year =. Household Inflation Expectations: An Overview of Recent Insights for Monetary Policy , institution =

  21. [29]

    and Malmendier, U

    D'Acunto, F. and Malmendier, U. and Ospina, J. and Weber, M. , year =. Exposure to Grocery Prices and Inflation Expectations , journal =

  22. [30]

    and Malmendier, U

    D'Acunto, F. and Malmendier, U. and Weber, M. , year =. What Do the Data Tell Us About Inflation Expectations? , booktitle =

  23. [31]

    and Weber, M

    D'Acunto, F. and Weber, M. , year =. Why Survey-Based Subjective Expectations Are Meaningful and Important , journal =

  24. [32]

    and Mikosch, H

    Dibiasi, A. and Mikosch, H. and Sarferaz, S. , year =. Uncertainty Shocks, Adjustment Costs, and Firm Beliefs: Evidence from a Representative Survey , journal =

  25. [33]

    and Pfajfar, D

    Ehrmann, M. and Pfajfar, D. and Santoro, E. , year =. Consumers' Attitudes and Their Inflation Expectations , journal =

  26. [34]

    and Wabitsch, A

    Ehrmann, M. and Wabitsch, A. , year =. Central bank communication with non-experts - A road to nowhere? , journal =

  27. [35]

    and Spiegler, R

    Eliaz, K. and Spiegler, R. , year =. A Model of Competing Narratives , journal =

  28. [36]

    Ericsson, K. A. and Hoffman, R. R. and Kozbelt, A. and Williams, A. M. , year =. The Cambridge Handbook of Expertise and Expert Performance , doi =

  29. [37]

    and Zafar, B

    Fuster, A. and Zafar, B. , year =. Survey experiments on economic expectations , booktitle =

  30. [38]

    and Xiong, Y

    Gao, Y. and Xiong, Y. and Gao, X. and Jia, K. and Pan, J. and Bi, Y. and Dai, Y. and Sun, J. and Wang, M. and Wang, H. , year =. Retrieval-Augmented Generation for Large Language Models: A Survey , doi =

  31. [39]

    and Chan, X

    Ge, T. and Chan, X. and Wang, X. and Yu, D. and Mi, H. and Yu, D. , year =. Scaling Synthetic Data Creation with 1,000,000,000 Personas , doi =

  32. [40]

    and Haan, P

    Gohl, N. and Haan, P. and Michelsen, C. and Weinhardt, F. , year =. House price expectations , journal =

  33. [41]

    Marketing Science , volume=

    Frontiers: Can Large Language Models Capture Human Preferences? , author=. Marketing Science , volume=. 2024 , publisher=. doi:10.1287/mksc.2023.0306 , url=

  34. [42]

    and Dahl, G

    Gordon, R. and Dahl, G. B. , year =. Views among Economists: Professional Consensus or Point-Counterpoint? , journal =

  35. [43]

    and Pham, T

    Gorodnichenko, Y. and Pham, T. and Talavera, O. , year =. Central bank communication on social media: What, to whom, and how? , journal =

  36. [44]

    Journal of Economic Literature , Volume =

    Haaland, Ingar and Roth, Christopher and Stantcheva, Stefanie and Wohlfart, Johannes , Title =. Journal of Economic Literature , Volume =. 2025 , Month =. doi:10.1257/jel.20251780 , URL =

  37. [45]

    and Roth, C

    Haaland, I. and Roth, C. and Wohlfart, J. , year =. Designing Information Provision Experiments , journal =

  38. [46]

    , year =

    Hagendorff, T. , year =. Deception abilities emerged in large language models , journal =

  39. [47]

    and Hangartner, D

    Hainmueller, J. and Hangartner, D. and Yamamoto, T. , year =. Validating vignette and conjoint survey experiments against real-world behavior , journal =

  40. [48]

    , year =

    Halterman, A. , year =. Synthetically generated text for supervised text analysis , journal =

  41. [49]

    Hansen, A. L. and Horton, J. J. and Kazinnik, S. and Puzzello, D. and Zarifhonarvar, A. , year =. Simulating the Survey of Professional Forecasters , journal =

  42. [50]

    2023 , institution=

    Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? , author=. 2023 , institution=

  43. [51]

    , year =

    Jonung, L. , year =. Perceived and Expected Rates of Inflation in Sweden , journal =

  44. [52]

    Lamla, M. J. and Maag, T. , year =. The Role of Media for Inflation Forecast Disagreement of Households and Professional Forecasters , journal =

  45. [53]

    Li, Nian and Gao, Chen and Li, Mingyu and Li, Yong and Liao, Qingmin , booktitle=

  46. [54]

    and Wu, J

    Ouyang, L. and Wu, J. and Jiang, X. and Almeida, D. and Wainwright, C. L. and Mishkin, P. and Zhang, C. and Agarwal, S. and Slama, K. and Ray, A. and Schulman, J. and Hilton, J. and Kelton, F. and Miller, L. and Simens, M. and Askell, A. and Welinder, P. and Christiano, P. and...

  47. [55]

    and Yun, H

    Ouyang, S. and Yun, H. and Zheng, X. , year =. How Ethical Should. doi:10.48550/ARXIV.2406.01168 , howpublished =

  48. [56]

    , year =

    Pearl, J. , year =. Causality: Models, Reasoning, and Inference , doi =

  49. [57]

    Sloman, S. A. and Lagnado, D. , year =. Causality in Thought , journal =

  50. [58]

    Souleles, N. S. , year =. Expectations, Heterogeneous Forecast Errors, and Consumption: Micro Evidence from the Michigan Consumer Sentiment Surveys , journal =

  51. [59]

    , year =

    Spiegler, R. , year =. Bayesian Networks and Boundedly Rational Expectations , journal =

  52. [60]

    , year =

    Spiegler, R. , year =. Behavioral Implications of Causal Misperceptions , journal =

  53. [61]

    and Brenninkmeijer, C.-F

    Tranchero, M. and Brenninkmeijer, C.-F. and Murugan, A. and Nagaraj, A. , year =. Theorizing with Large Language Models , institution =

  54. [62]

    and Kahneman, D

    Tversky, A. and Kahneman, D. , year =. Availability: A heuristic for judging frequency and probability , journal =

  55. [63]

    and Morgenstern, J

    Wang, A. and Morgenstern, J. and Dickerson, J. P. , year =. Large language models that replace human participants can harmfully misportray and flatten identity groups , journal =

  56. [64]

    and D'Acunto, F

    Weber, M. and D'Acunto, F. and Gorodnichenko, Y. and Coibion, O. , year =. The Subjective Inflation Expectations of Households and Firms: Measurement, Determinants, and Implications , journal =

  57. [65]

    Wu, J. C. and Xi, J. and Xie, S. , year =. LLM Survey Framework: Coverage, Reasoning, Dynamics, Identification , institution =

  58. [66]

    and Li, A

    Yang, A. and Li, A. and Yang, B. and Zhang, B. and Hui, B. and Zheng, B. and Yu, B. and Gao, C. and Huang, C. and Lv, C. and Zheng, C. and Liu, D. and Zhou, F. and Huang, F. and Hu, F. and Ge, H. and Wei, H. and Lin, H. and Tang, J. and Qiu, Z. , year =. Qwen3 Technical Report , doi =

  59. [67]

    Advances in Neural Information Processing Systems , volume=

    Large language model as attributed training data generator: A tale of diversity and bias , author=. Advances in Neural Information Processing Systems , volume=

  60. [68]

    , year =

    Zarifhonarvar, A. , year =. Generating inflation expectations with large language models , journal =

  61. [69]

    Zhao, W. X. and Zhou, K. and Li, J. and Tang, T. and Wang, X. and Hou, Y. and Min, Y. and Zhang, B. and Zhang, J. and Dong, Z. and Du, Y. and Yang, C. and Chen, Y. and Chen, Z. and Jiang, J. and Ren, R. and Li, Y. and Tang, X. and Liu, Z. and Wen, J.-R. , year =. A Survey of L...

  62. [70]

    Review of Economics and Statistics , volume=

    The price is right: Updating inflation expectations in a randomized price information experiment , author=. Review of Economics and Statistics , volume=. 2016 , publisher=

  63. [71]

    2026 , publisher=

    Expectations matter: The new causal macroeconomics of surveys and experiments , author=. 2026 , publisher=

  64. [72]

    Journal of Economic Perspectives , volume=

    The billion prices project: Using online prices for measurement and research , author=. Journal of Economic Perspectives , volume=. 2016 , publisher=

  65. [73]

    Journal of Economic Perspectives , volume=

    Household surveys in crisis , author=. Journal of Economic Perspectives , volume=. 2015 , publisher=

  66. [74]

    Journal of Economic Perspectives , volume=

    Evolving measurement for an evolving economy: thoughts on 21st century US economic statistics , author=. Journal of Economic Perspectives , volume=. 2019 , publisher=

  67. [75]

    Humanities and Social Sciences Communications , volume=

    Large language models empowered agent-based modeling and simulation: A survey and perspectives , author=. Humanities and Social Sciences Communications , volume=. 2024 , publisher=

  68. [76]

    2026 , institution=

    Information Treatments, Hypotheticals, and Event Studies: Comparative Estimates , author=. 2026 , institution=

  69. [77]

    Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , pages=

    Generative agents: Interactive simulacra of human behavior , author=. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , pages=

  70. [78]

    2026 , publisher =

    Mou, Xinyi and Ding, Xuanwen and He, Qi and Wang, Liang and Liang, Jingcong and Zhang, Xinnong and Sun, Libo and Lin, Jiayu and Zhou, Jie and Xuanjing, Huang and Wei, Zhongyu , title =. 2026 , publisher =. doi:10.1145/3800683 , journal =

  71. [79]

    Bank Run, Interrupted: Modeling Deposit Withdrawals with Generative

    Kazinnik, Sophia , journal=. Bank Run, Interrupted: Modeling Deposit Withdrawals with Generative. 2026 , note=

  72. [80]

    Jinghua Piao and Yuwei Yan and Jun Zhang and Nian Li and Junbo Yan and Xiaochong Lan and Zhihong Lu and Zhiheng Zheng and Jing Yi Wang and Di Zhou and Chen Gao and Fengli Xu and Fang Zhang and Ke Rong and Jun Su and Yong Li , year =

  73. [81]

    2024 , publisher=

    Meng, Juanjuan , journal=. 2024 , publisher=

  74. [82]

    Evaluating the statistical realism of

    Xie, Yueqi and Liang, Lemeng and Li, Shuzhen and Lu, Yifu and Xiao, Zhiwen and Shi, Mengdi and Huang, Junming and Wang, Mengdi and Xie, Yu , journal=. Evaluating the statistical realism of. 2026 , doi=

  75. [83]

    and Dreber, Anna and Holzmeister, Felix and Ho, Teck-Hua and Huber, J

    Camerer, Colin F. and Dreber, Anna and Holzmeister, Felix and Ho, Teck-Hua and Huber, J. Evaluating the replicability of social science experiments in. Nature Human Behaviour , volume=. 2018 , publisher=

  76. [84]

    Camerer, Colin F. and Dreber, Anna and Forsell, Eskil and Ho, Teck-Hua and Huber, Jürgen and Johannesson, Magnus and Kirchler, Michael and Almenberg, Johan and Altmejd, Adam and Chan, Taizan and Heikensten, Emma and Holzmeister, Felix and Imai, Taisuke and Isaksson, Siri and N...

  77. [85]

    and Arriaga, Rosa I

    Aher, Gati V. and Arriaga, Rosa I. and Kalai, Adam Tauman , year =. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies , booktitle =

  78. [86]

    Journal of Monetary Economics , volume=

    Implications of rational inattention , author=. Journal of Monetary Economics , volume=. 2003 , publisher=

  79. [87]

    Computers, Communications, and the Public Interest , pages =

    Designing Organizations for an Information-Rich World , author =. Computers, Communications, and the Public Interest , pages =. 1971 , publisher =

  80. [88]

    Journal of Economic Literature , Volume =

    Loewenstein, George and Wojtowicz, Zachary , Title =. Journal of Economic Literature , Volume =. 2025 , Month =. doi:10.1257/jel.20241665 , URL =

  81. [89]

    Mei, Qiaozhu and Xie, Yutong and Yuan, Walter and Jackson, Matthew O , journal=. A. 2024 , publisher=

  82. [90]

    Handbook of Economic Expectations , pages=

    Bayesian learning , author=. Handbook of Economic Expectations , pages=. 2023 , publisher=

  83. [91]

    The Journal of Finance , volume=

    Investor psychology and security market under-and overreactions , author=. The Journal of Finance , volume=. 1998 , publisher=

  84. [92]

    Do you know that

    Coibion, Olivier and Gorodnichenko, Yuriy and Kumar, Saten and Ryngaert, Jane , journal=. Do you know that. 2021 , publisher=

  85. [93]

    Review of Economics and Statistics , volume=

    Forecaster (mis-) behavior , author=. Review of Economics and Statistics , volume=. 2024 , publisher=

  86. [94]

    Position:

    Anthis, Jacy Reese and Liu, Ryan and Richardson, Sean M and Kozlowski, Austin C and Koch, Bernard and Brynjolfsson, Erik and Evans, James and Bernstein, Michael S , booktitle=. Position:

  87. [95]

    Findings of the Association for Computational Linguistics: ACL 2025 , pages=

    Mixture-of-personas language models for population simulation , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=

  88. [96]

    Cecere, Nicola and Bacciu, Andrea and Fern. Monte. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025) , pages=

  89. [97]

    2018 , edition=

    Econometric Analysis , author=. 2018 , edition=

  90. [98]

    and Thomas, Joy A

    Cover, Thomas M. and Thomas, Joy A. , title =

  91. [99]

    2011 , publisher=

    Econometrics , author=. 2011 , publisher=

  92. [100]

    The Review of Economic Studies , volume=

    Home price expectations and behaviour: Evidence from a randomized information experiment , author=. The Review of Economic Studies , volume=. 2019 , publisher=

  93. [101]

    Annual Review of Economics , volume=

    Large Language Models: An Applied Econometric Framework , author=. Annual Review of Economics , volume=. 2026 , publisher=. doi:10.1146/annurev-economics-120925-105620 , url=

  94. [102]

    Nature , volume =

    Spirling, Arthur , title =. Nature , volume =. 2023 , doi =

  95. [103]

    Harvard Data Science Review , volume =

    Chen, Lingjiao and Zaharia, Matei and Zou, James , title =. Harvard Data Science Review , volume =. 2024 , doi =

  96. [104]

    Journal of Monetary Economics , volume=

    People's understanding of inflation , author=. Journal of Monetary Economics , volume=. 2024 , publisher=

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.