Pith. sign in

REVIEW 4 cited by

Mining Causality: AI-Assisted Search for Instrumental Variables

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.14202 v3 pith:NJUFCBA4 submitted 2024-09-21 econ.EM stat.APstat.MEstat.ML

classification econ.EMstat.APstat.MEstat.ML
keywords searchvariablesfindinginstrumentallanguagelargellmsmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process, and justifying its validity -- especially exclusion restrictions -- is largely rhetorical. We propose using large language models (LLMs) to search for new IVs through narratives and counterfactual reasoning, similar to how a human researcher would. The stark difference, however, is that LLMs can dramatically accelerate this process and explore an extremely large search space. We demonstrate how to construct prompts to search for potentially valid IVs. We contend that multi-step and role-playing prompting strategies are effective for simulating the endogenous decision-making processes of economic agents and for navigating language models through the realm of real-world scenarios, rather than anchoring them within the narrow realm of academic discourses on IVs. We apply our method to three well-known examples in economics: returns to schooling, supply and demand, and peer effects. We then extend our strategy to finding (i) control variables in regression and difference-in-differences and (ii) running variables in regression discontinuity designs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

    stat.ML 2026-07 conditional novelty 7.0 of 10

    CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.

  2. Do LLMs Act as Repositories of Causal Knowledge?

    econ.EM 2024-12 conditional novelty 5.0 of 10

    Commercial LLMs agree only modestly with expert confounder lists for the Coronary Drug Project and flip answers under trivial prompt changes, so they cannot yet act as reliable causal knowledge repositories.

  3. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

  4. Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention

    cs.AI 2024-11 reject novelty 2.0 of 10

    Synthetic data intervention is reported to reduce sycophancy in GPT-4o on 100 true-false questions, but the experiment does not establish that the model was actually trained.

Pith tools