REVIEW 20 cited by
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The causal capabilities of large language models (LLMs) are a matter of significant debate, with critical implications for the use of LLMs in societally impactful domains such as medicine, science, law, and policy. We conduct a "behavorial" study of LLMs to benchmark their capability in generating causal arguments. Across a wide range of tasks, we find that LLMs can generate text corresponding to correct causal arguments with high probability, surpassing the best-performing existing methods. Algorithms based on GPT-3.5 and 4 outperform existing algorithms on a pairwise causal discovery task (97%, 13 points gain), counterfactual reasoning task (92%, 20 points gain) and event causality (86% accuracy in determining necessary and sufficient causes in vignettes). We perform robustness checks across tasks and show that the capabilities cannot be explained by dataset memorization alone, especially since LLMs generalize to novel datasets that were created after the training cutoff date. That said, LLMs exhibit unpredictable failure modes, and we discuss the kinds of errors that may be improved and what are the fundamental limits of LLM-based answers. Overall, by operating on the text metadata, LLMs bring capabilities so far understood to be restricted to humans, such as using collected knowledge to generate causal graphs or identifying background causal context from natural language. As a result, LLMs may be used by human domain experts to save effort in setting up a causal analysis, one of the biggest impediments to the widespread adoption of causal methods. Given that LLMs ignore the actual data, our results also point to a fruitful research direction of developing algorithms that combine LLMs with existing causal techniques. Code and datasets are available at https://github.com/py-why/pywhy-llm.
Forward citations
Cited by 20 Pith papers
-
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
In controlled synthetic worlds, more interventional pretraining grows the magnitude but not the sign of a language model's causal response; the observational content in the inference context decides the sign.
-
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.
-
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
Intuitiveness of policy findings dominates LLM counterfactual accuracy, with chain-of-thought providing almost no benefit on counter-intuitive cases and familiarity with citations unrelated to performance.
-
Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
By combining topological ordering, pruning, and disentanglement, CREATOR identifies linearly mixed latent causal variables up to permutation-and-scale ambiguity using only non-Gaussian noise.
-
Decomposed Entailment for Factuality Checking and Hallucination Detection
HallDetect detects source-grounded hallucinations by decomposing responses into atomic claims and verifying each with a compact NLI model over multi-scale source chunks, outperforming frugal generative baselines on th...
-
Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation
Adding LLM-generated, uncertainty-targeted semantic representations — split into assignment and heterogeneity channels and routed asymmetrically — improves finite-sample CATE estimates for most of ten neural host lear...
-
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
CaM-Wolf is a multimodal Werewolf agent that perceives player video, reasons about hidden roles with a counterfactual-intervention-trained RL reasoner, and responds through an animated avatar.
-
EviDAG: Auditable Causal DAG Authoring with Biomedical Literature
EviDAG converts free-text study variables into an auditable, constraint-checked causal DAG whose edges carry verified, quoted evidence from a frozen PubMed snapshot.
-
Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market
Using FOMC minutes to propose regime-shift candidates and a lenient text check to ratify data-detected candidates, the pipeline reaches F1=0.82 on 26 monetary-policy anchors, beating every data-only baseline.
-
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
A new benchmark and training strategy show LLMs trained to internalize causal reasoning steps are less fooled by semantically similar, label-flipped questions than models using explicit chain-of-thought.
-
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks
A causality-inspired LLM planning framework with task graphs, counterfactual rule checks, and busy-rate path assignment improves multi-agent Minecraft task completion in reported experiments.
-
Math Natural Language Inference: this should be easy!
A new Math NLI corpus from category theory abstracts plus a ten-model evaluation shows LLM unanimous votes approach human labels (88%) but individual LLMs still make basic math reasoning errors.
-
Zero-Shot Event Causality Identification via Multi-source Evidence Fuzzy Aggregation with Large Language Models
MEFA aggregates probability outputs from six causality sub-tasks via a fuzzy Choquet integral, improving zero-shot event causality identification by 6.2% F1 over the best unsupervised baseline.
-
LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery
LLMs are unreliable causal reasoners, so they should be limited to non-decisional search support in causal discovery algorithms.
-
ACCESS : A Benchmark for Abstract Causal Event Discovery and Reasoning
ACCESS is a new benchmark of 725 abstract event clusters and 1,494 causal relations from GLUCOSE, with experiments showing that LLMs struggle at abstraction and causal discovery but improve on QA when the correct caus...
-
Ethical Considerations of Large Language Models in Game Playing
In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.
-
Do Large Language Models Reason Causally Like Us? Even Better?
GPT-4o, Gemini-Pro, and Claude show less associative bias than human reasoners on collider graphs, making them more normatively aligned in likelihood judgments, but none fully demonstrates explaining away.
-
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
A risk score built from confidence and vote entropy is proposed for triaging LLM qualitative coding, but the central R-squared=0.979 result is inflated by the definitions of agreement and diversity.
-
GraphRAG-Causal: A novel graph-augmented framework for causal reasoning and annotation in news
A graph-retrieval-augmented LLM pipeline for causal news classification reports 82.1% F1 with 20 examples, but likely leaks test data into its retrieval store.
-
Causal MAS: A Survey of Large Language Model Architectures for Discovery and Effect Estimation
A structured survey that defines and catalogs multi-agent LLM systems for causal reasoning, discovery, and effect estimation, including their architectures, benchmarks, and applications.
Discussion (0). Continue with ORCID to comment.