Pith. sign in

REVIEW 7 cited by

Causal Parrots: Large Language Models May Talk Causality But Are Not Causal

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.13067 v1 pith:2DAFY5AS submitted 2023-08-24 cs.AI cs.CL

classification cs.AIcs.CL
keywords causallanguagellmsmodelsparrotsdataevenfacts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Some argue scale is all what is needed to achieve AI, covering even causal models. We make it clear that large language models (LLMs) cannot be causal and give reason onto why sometimes we might feel otherwise. To this end, we define and exemplify a new subgroup of Structural Causal Model (SCM) that we call meta SCM which encode causal facts about other SCM within their variables. We conjecture that in the cases where LLM succeed in doing causal inference, underlying was a respective meta SCM that exposed correlations between causal facts in natural language on whose data the LLM was ultimately trained. If our hypothesis holds true, then this would imply that LLMs are like parrots in that they simply recite the causal knowledge embedded in the data. Our empirical analysis provides favoring evidence that current LLMs are even weak `causal parrots.'

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

    stat.ML 2026-07 conditional novelty 7.0 of 10

    CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.

  2. STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Frontier LLM agents detect hidden supply-chain stress almost equally well (84–88% of episodes) but vary from skill 0.62 to −0.23, with two of four models acting worse than ignoring symptoms.

  3. Decomposed Entailment for Factuality Checking and Hallucination Detection

    cs.CL 2026-08 conditional novelty 6.0 of 10

    HallDetect detects source-grounded hallucinations by decomposing responses into atomic claims and verifying each with a compact NLI model over multi-scale source chunks, outperforming frugal generative baselines on th...

  4. Causal-Invariant Cross-Domain Out-of-Distribution Recommendation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    CICDOR learns two causal DAGs for shared and domain-specific user preferences, uses an LLM guided by the FCI algorithm to extract confounders from reviews, and reports consistent accuracy gains over twelve baselines o...

  5. Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning a 3B LLM on randomly symbolized reasoning questions reduces spurious-correlation failures and improves OOD accuracy on CLadder and PrOntoQA.

  6. Do Large Language Models Reason Causally Like Us? Even Better?

    cs.AI 2025-02 conditional novelty 5.0 of 10

    GPT-4o, Gemini-Pro, and Claude show less associative bias than human reasoners on collider graphs, making them more normatively aligned in likelihood judgments, but none fully demonstrates explaining away.

  7. GraphRAG-Causal: A novel graph-augmented framework for causal reasoning and annotation in news

    cs.IR 2025-06 reject novelty 4.0 of 10

    A graph-retrieval-augmented LLM pipeline for causal news classification reports 82.1% F1 with 20 examples, but likely leaks test data into its retrieval store.

Pith tools