REVIEW 5 major objections 5 minor 1 cited by
For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read GPT-4 treats backgrounded content as islands, exactly as humans do.
desk verdict Solid correlational evidence that GPT-4 tracks information-structure-to-acceptability mappings; the causal Study 2 is confounded by an explicit instruction and post-hoc item selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the BCI principle—'Backgrounded Constituents are Islands'—which states that a constituent's unavailability for long-distance dependencies scales with how backgrounded its content is. The paper operationalizes backgroundedness with the negation task: after a negated sentence, a question about the key constituent is answered 'probably yes' if that content survives negation, meaning it is presupposed or backgrounded rather than at-issue. This gradient backgroundedness score, collected on canonical base sentences, is then correlated with independently elicited acceptability ratings on LDDs built from those same sentences. In Study 2 the machinery becomes causal: a context sentence with lexical emphasis on the to-be-queried constituent increases that constituent's prominence and, in turn, GPT-4's acceptability rating of the subsequent wh-question.
What would settle it
Compare GPT-4's backgroundedness ratings on items where human presupposition judgments diverge from logical entailment (for example, negated clausal complements whose content is mentioned but not presupposed); if backgroundedness no longer predicts LDD acceptability on those items, the claimed role of information structure is not supported. Replacing the negation task with a focus-based measure, such as question-answer congruence, and checking whether the LDD predictability survives would provide a second decisive test.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that GPT-4's explicit metalinguistic judgments about information structure predict its independent acceptability ratings on long-distance dependency constructions, in the direction predicted by the Backgrounded Constituents are Islands (BCI) principle: constituents rated more backgrounded in a base sentence receive lower acceptability ratings when they are extracted in LDDs. This holds for wh-questions, discourse-linked questions, relative clauses, and it-clefts, and it holds with newly generated stimuli designed to rule out training-data contamination. Study 2 goes further: when a preceding context sentence emphasizes the to-be-queried constituent, GPT-4 gives higher acceptability ratings to the following wh-question than when the same context is presented without emphasis. The paper concludes that GPT-4 captures a systematic, causal relationship between information structure and syntactic acceptability that parallels human psycholinguistic results.
Load-bearing premise
The load-bearing assumption is that GPT-4's 'probably yes' answers on the negation task reflect genuine presupposition or backgroundedness, rather than a shallow inference that a negated sentence still mentions the content; if the latter were true, the correlation with LDD acceptability could be driven by an unmeasured third factor.
Editorial extensions
If this is right
- If the pattern holds, GPT-4 can serve as a zero-shot informant for gradient acceptability judgments, letting researchers probe island constraints and information structure without large human norming samples.
- The BCI principle gains a new form of evidence: a model with no explicit linguistic rules reproduces the interaction, strengthening discourse-functional accounts over purely syntactic ones.
- Because Study 1b used fresh stimuli and added it-clefts, the effect transfers to a construction not present in the original human data, arguing against contamination by memorized examples.
- The causal emphasis effect from Study 2 implies that acceptability judgments in GPT-4 are context-sensitive, not static ratings of sentence form alone.
Reading between the lines
- A testable extension: if the BCI interaction is tied to scale, smaller or open-weight models should show a weaker or absent effect; a graded pattern across model sizes would clarify whether this is an emergent property of language-model training objectives.
- The same prominence manipulation could be probed in generation: GPT-4 may be more likely to produce wh-questions that extract constituents made prominent in the preceding context, offering a behavioral test beyond explicit ratings.
- The negation-task assumption can be stress-tested by comparing GPT-4's ratings with formal presupposition-projection diagnostics; divergence would mark the boundary of the model's information-structure sensitivity.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper tests whether GPT-4, probed with zero-shot explicit metalinguistic tasks, replicates the human pattern whereby backgroundedness of a constituent in a base sentence predicts acceptability of long-distance dependency (LDD) constructions. In Study 1a, GPT-4's backgroundedness judgments (negation task) and acceptability ratings on 144 human stimuli show a significant interaction: increased backgroundedness predicts lower acceptability for LDDs more than for base sentences. Study 1b replicates with new stimuli and an added construction. Study 2 presents context sentences with or without emphasis on the to-be-queried constituent and reports that emphasis raises GPT-4's acceptability ratings on subsequent wh-questions, interpreted as a causal effect of information structure. The abstract and General Discussion conclude that GPT-4 exhibits an emergent understanding of information structure's role in syntactic acceptability.
Significance. If the causal claim held, this would be a significant demonstration that a large language model captures a subtle form-function interaction, with implications for both LM evaluation and theories of island constraints. The paper's strengths include zero-shot probing with temperature zero, ten repetitions per stimulus, replication on fresh stimuli to address contamination, public data, and a direct attempt at causal manipulation. However, the causal conclusion is currently undermined by confounds in Study 2 and by selective item analysis, so the central contribution remains the correlational finding. The correlational studies are informative and, with proper reporting of human correlations and model details, could constitute a solid contribution, but the paper's strongest claim ('confirms a causal relationship') is not yet supported.
major comments (5)
- [Study 2, Table 4] The Emphasis condition differs from the No-emphasis condition not only in ALL-CAPS emphasis but also in the explicit instruction 'Please focus on the part in ALL CAPS.' This instruction is a demand characteristic that could raise ratings for the queried constituent independently of any representation of information structure. The reported causal claim therefore is not supported. The authors should add a control condition that includes an analogous instruction (e.g., 'Please focus on the part in lowercase' or 'Please focus on the whole sentence') or orthogonally vary emphasis and instruction in a 2x2 design.
- [Study 2, Results] The analysis is restricted to the 72 items with non-emphasis mean acceptability below 6. This is a post hoc selection on the control-condition outcome, and the reported effect (ß = 0.44, t = 8.48) may partly reflect regression to the mean. Please analyze the full set of items as the primary analysis, justify any exclusion criterion beforehand, and report the effect size for the full set; if the effect only appears in the selected subset, state that clearly.
- [Study 1a, Results/Introduction] The manuscript claims 'strong correlations with human judgments on the same stimuli' but no correlation coefficient, test statistic, or confidence interval is reported anywhere. Because the paper's framing compares GPT-4 with human BCI results, this correlation is essential. Please report the correlation between GPT-4 and human mean ratings (per item) for both backgroundedness and acceptability, and clarify whether the interaction slopes are quantitatively similar.
- [Methods (proprietary model)] The exact model version is not specified (only 'GPT-4 API, Jan 2025'), and the random-effect structure is described only as 'the maximal random effect structure convergence allowed.' This hampers reproducibility. Please report the exact model identifier (e.g., gpt-4-0613 or gpt-4-turbo-2025-01) and the full model formula, including the random effects used in the ordinal model.
- [Study 1, Backgroundedness task] The negation task is assumed to measure backgroundedness in GPT-4 just as it does in humans. Given that GPT-4's 'probably yes' answers might reflect shallow lexical or logical inferences (e.g., that a negated event still mentions the content), the paper should provide convergent validity evidence for this measure, for example by showing that the same backgroundedness judgments correlate with an independent information-structure probe or with the emphasis manipulation in Study 2 once the confound is removed.
minor comments (5)
- [Introduction] The phrase 'Here was ask' should be 'Here we ask.'
- [Table 2] The scale label '4 means probably yet' should be '4 means probably yes.'
- [General Discussion] The text says 'three types of long-distance dependency constructions' and lists wh-Qs, discourse-linked questions, and relative clauses, but the Methods of Study 1a state that four types (including it-clefts) were collected; please reconcile the count and clarify which constructions were included in each study.
- [Study 1a] The reference 'Authors (2023)' in the description of the human study should be replaced with the proper citation 'Cuneo and Goldberg (2023).'
- [Figure 1] Figure 1 is described in the text but not visible in the manuscript; please ensure the figure is included and legible, as it is the only visual comparison of human and GPT-4 data.
Circularity Check
No significant circularity: GPT-4 judgments are new empirical probes; self-citations supply the hypothesis, not the derived outcome.
full rationale
The paper claims GPT-4's backgroundedness judgments predict its own independent acceptability ratings on LDDs (Studies 1a/1b) and that emphasis in a context sentence raises acceptability (Study 2). These claims are empirical correlations and effects, not identities. Backgroundedness and acceptability are elicited in separate zero-shot prompts (Tables 1 and 2), with no fitted parameters, no calibration, and no shared response equation; the ordinal model only estimates the relationship, so the interaction is not constructed by the analysis. The BCI hypothesis is imported from the authors' prior human work, but the prior results do not entail what GPT-4 will do; the current data are new and could have failed. Self-citations to Cuneo and Goldberg (2023), Goldberg (2006, 2013), and Namboodiripad et al. (2022) supply the theoretical motivation, not the GPT-4 evidence, and independent labs are also cited for the human phenomenon (e.g., Lu, Pan, and Degen 2024; Winckel et al. 2025). Study 1b uses novel stimuli and an added construction, mitigating contamination. The main validity concern is Study 2's manipulation: the Emphasis condition adds an explicit 'Please focus on the part in ALL CAPS' instruction together with typographic emphasis, and the analysis selects items with non-emphasis ratings below 6; these can inflate or create the effect through demand characteristics and regression to the mean. However, that is a confound in causal inference, not a circular derivation in which a prediction reduces to its input by definition. No step meets the quoted-equation/reduction bar.
Assumptions & free parameters
assumptions (3)
- domain assumption The negation task (e.g., whether a constituent survives under sentence negation) is a valid operationalization of backgroundedness in GPT-4.
- domain assumption Zero-shot explicit Likert ratings by GPT-4 are a valid measure of acceptability, despite critiques (Hu and Levy 2023) that prompting is not equivalent to log-probability measures.
- domain assumption The GPT-4 API version used (as of Jan 2025) behaves consistently across queries at temperature 0.
Cite this review
Pith. "Pith review of For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies." pith.science (2026). https://pith.science/paper/HTOGAAWV
@misc{pith2026250509005,
author = {Pith},
title = {Pith review of: For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies},
year = {2026},
howpublished = {\url{https://pith.science/paper/HTOGAAWV}},
note = {Machine review of arXiv:2505.09005}
}
read the original abstract
It remains debated how well any LM understands natural language or generates reliable metalinguistic judgments. Moreover, relatively little work has demonstrated that LMs can represent and respect subtle relationships between form and function proposed by linguists. We here focus on a particular such relationship established in recent work: English speakers' judgments about the information structure of canonical sentences predicts independently collected acceptability ratings on corresponding 'long distance dependency' [LDD] constructions, across a wide array of base constructions and multiple types of LDDs. To determine whether any LM captures this relationship, we probe GPT-4 on the same tasks used with humans and new extensions.Results reveal reliable metalinguistic skill on the information structure and acceptability tasks, replicating a striking interaction between the two, despite the zero-shot, explicit nature of the tasks, and little to no chance of contamination [Studies 1a, 1b]. Study 2 manipulates the information structure of base sentences and confirms a causal relationship: increasing the prominence of a constituent in a context sentence increases the subsequent acceptability ratings on an LDD construction. The findings suggest a tight relationship between natural and GPT-4 generated English, and between information structure and syntax, which begs for further exploration.
Figures
Forward citations
Cited by 1 Pith paper
-
When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
VLMs use Hungarian word order to mark Topic and Focus but drastically underproduce the variable strategies humans show under conflicting discourse pressures, resembling mode collapse.
Reference graph
Works this paper leans on
-
[2]
Proceedings of the Cognitive Science Society. and Degen, 2024; Namboodiripad et al., 2022). In wh-Qs like (2 ) - (3) the wh-word (what) is prominent, while the causally related adjunct phrase “by admitting” is less backgrounded than the temporal adverb, “while admitting.” The claim is essentially that anomaly results when a speaker chooses to foreground a...
work page 2022
-
[3]
Additionally, GPT-4's understanding of information structure may be superficial or unreliable
Proceedings of the Cognitive Science Society. Additionally, GPT-4's understanding of information structure may be superficial or unreliable. Indeed, Sieker and Zarrieß (2023) find that BERT models perform poorly on certain types of implicatures related to a speaker choosing one word or construction rather than another (see also Jeretic et al. 2020). Srava...
work page 2023
-
[4]
Please focus on the part in ALL CAPS,
Proceedings of the Cognitive Science Society. Figure 1: Comparison between human data [Left panel]; GPT-4 on the same stimuli [middle]; GPT-4 on newly created stimuli [Right]. Raw backgroundedness judgments on base sentences (x-axis) and Acceptability ratings on each LDD (shades of red) and base sentences (blue) (y-axis). Data available here: https://rese...
work page 2024
-
[5]
striking interaction between these two factors only recently confirmed in English-speakers
Proceedings of the Cognitive Science Society. striking interaction between these two factors only recently confirmed in English-speakers. Specifically, Study 1a replicates a study on humans (Cuneo and Goldberg, 2023) and establishes that GPT-4 reliably assesses information structure on canonical (base) sentences; moreover, those judgments predict its inde...
work page 2023
-
[6]
Testing AI on Language Comprehension Tasks Reveals Insensitivity to Underlying Meaning
Proceedings of the Cognitive Science Society. Acknowledgements We are grateful to Arielle Belluck for editing a penultimate version of the manuscript, and to Princeton’s Natural and Artificial Minds program for funding. References Ambridge, B., & Goldberg, A. E. (2008). The island status of clausal complements: Evidence in favor of an information structur...
arXiv 2008
-
[7]
BLiMP: The Benchmark of Linguistic Minimal Pairs for English
Proceedings of the Cognitive Science Society. Potts, C. (2004). The logic of conventional implicatures. OUP. Potts, C. (2023). Characterizing English Preposing in PP constructions. LingBuzz lingbuzz/007495. Ross, J. R. (1967). Constraints on variables in syntax. Ph.D. Thesis, Massachusetts Institute of Technology. Sieker, J., & Zarrieß, S. (2023, December...
arXiv 2004
-
[2025]
Proceedings of the Cognitive Science Society. For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies Nicole Cuneo1, Eleanor Graves, Supantho Rakshit2, Adele E. Goldberg1 {nicole.cuneo@, eg5817@, r.supantho@, adele@}princeton.edu Departments of Psychology1 and ECE2 Princeton University, Princeton NJ 08544 USA A...
work page 2022
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.