REVIEW 7 cited by
Causal Parrots: Large Language Models May Talk Causality But Are Not Causal
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Some argue scale is all what is needed to achieve AI, covering even causal models. We make it clear that large language models (LLMs) cannot be causal and give reason onto why sometimes we might feel otherwise. To this end, we define and exemplify a new subgroup of Structural Causal Model (SCM) that we call meta SCM which encode causal facts about other SCM within their variables. We conjecture that in the cases where LLM succeed in doing causal inference, underlying was a respective meta SCM that exposed correlations between causal facts in natural language on whose data the LLM was ultimately trained. If our hypothesis holds true, then this would imply that LLMs are like parrots in that they simply recite the causal knowledge embedded in the data. Our empirical analysis provides favoring evidence that current LLMs are even weak `causal parrots.'
Forward citations
Cited by 7 Pith papers
-
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.
-
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
Frontier LLM agents detect hidden supply-chain stress almost equally well (84–88% of episodes) but vary from skill 0.62 to −0.23, with two of four models acting worse than ignoring symptoms.
-
Decomposed Entailment for Factuality Checking and Hallucination Detection
HallDetect detects source-grounded hallucinations by decomposing responses into atomic claims and verifying each with a compact NLI model over multi-scale source chunks, outperforming frugal generative baselines on th...
-
Causal-Invariant Cross-Domain Out-of-Distribution Recommendation
CICDOR learns two causal DAGs for shared and domain-specific user preferences, uses an LLM guided by the FCI algorithm to extract confounders from reviews, and reports consistent accuracy gains over twelve baselines o...
-
Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training
Fine-tuning a 3B LLM on randomly symbolized reasoning questions reduces spurious-correlation failures and improves OOD accuracy on CLadder and PrOntoQA.
-
Do Large Language Models Reason Causally Like Us? Even Better?
GPT-4o, Gemini-Pro, and Claude show less associative bias than human reasoners on collider graphs, making them more normatively aligned in likelihood judgments, but none fully demonstrates explaining away.
-
GraphRAG-Causal: A novel graph-augmented framework for causal reasoning and annotation in news
A graph-retrieval-augmented LLM pipeline for causal news classification reports 82.1% F1 with 20 examples, but likely leaks test data into its retrieval store.
Discussion (0). Continue with ORCID to comment.