{"id":"ebc9f1e4-5391-46f4-b557-cd1b2480698f","arxiv_id":"2505.09005","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4 reproduces the human interaction between backgroundedness and acceptability for long-distance dependencies, and its ratings increase when the queried constituent is emphasized in context.","lead":"GPT-4's judgments about which parts of a sentence are backgrounded predict its own acceptability ratings for long-distance dependency constructions, matching a pattern first documented in people. The result suggests large language models encode a subtle link between information structure and syntax, and it offers a new way to probe functional constraints in AI language behavior.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Study 2's causal claim is confounded: the Emphasis condition adds an explicit focus instruction alongside ALL-CAPS emphasis, and the analysis excludes items with high baseline acceptability.","rationale":"The reader's weakest assumption concerns the validity of the negation task in Study 1. That is a reasonable concern, but I think the more load-bearing issue is the causal inference in Study 2. The correlational Studies 1a/1b already control for response bias via the interaction with base sentences, so a third-variable account of the backgroundedness measure is less threatening. Study 2, however, is the only place where the paper claims causality, and its manipulation is confounded: the Emphasis condition pairs ALL-CAPS lexical stress with an explicit instruction to focus on the caps, while the No-emphasis condition has no analogous instruction. This makes the result compatible with simple instruction-following. The additional post hoc restriction to items with baseline mean <6 compounds the problem, since selection on the control outcome can produce regression artifacts. The paper's headline claim — that increasing prominence causes higher acceptability — therefore rests on the least secure part of the design. I would keep the CONDITIONAL verdict because the correlational findings are still valuable and the causal issue is fixable with a control experiment; the concern is not grounds for rejection, but it must be addressed before the causal claim can be accepted.","tokens_in":8350,"tokens_out":11828,"duration_ms":120370,"concrete_test":"Run a 2x2 experiment crossing emphasis (ALL CAPS vs plain) with instruction (the current explicit 'focus on ALL CAPS' prompt vs a neutral instruction such as 'Please read the context sentence carefully'), and analyze all 144 items without baseline-based exclusion. The causal claim is supported only if ALL CAPS raises acceptability ratings in the neutral-instruction condition; otherwise the reported effect is an artifact of instruction-following rather than information structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's only causal evidence comes from Study 2, but the Emphasis condition differs from No-emphasis in two ways: the context sentence uses ALL CAPS, and the prompt explicitly instructs GPT-4 to 'Please focus on the part in ALL CAPS' (Table 4). The No-emphasis condition has no analogous instruction. This explicit meta-linguistic cue is a demand characteristic: telling the model to focus on a constituent can raise its acceptability rating for a wh-question about that constituent regardless of any representation of information structure. In addition, the reported analysis is restricted to the 72 items with non-emphasis mean acceptability <6, a post hoc selection on the control-condition outcome. Such selection can create a spurious emphasis effect via regression to the mean. Because the abstract's 'causal relationship' claim and the General Discussion's 'emergent understanding' inference rest on Study 2, the central claim is not yet supported. The correlational Studies 1a/1b remain informative, but they cannot by themselves establish the causal role of information structure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper tests whether GPT-4, probed with zero-shot explicit metalinguistic tasks, replicates the human pattern whereby backgroundedness of a constituent in a base sentence predicts acceptability of long-distance dependency (LDD) constructions. In Study 1a, GPT-4's backgroundedness judgments (negation task) and acceptability ratings on 144 human stimuli show a significant interaction: increased backgroundedness predicts lower acceptability for LDDs more than for base sentences. Study 1b replicates with new stimuli and an added construction. Study 2 presents context sentences with or without emphasis on the to-be-queried constituent and reports that emphasis raises GPT-4's acceptability ratings on subsequent wh-questions, interpreted as a causal effect of information structure. The abstract and General Discussion conclude that GPT-4 exhibits an emergent understanding of information structure's role in syntactic acceptability.","tokens_in":8469,"tokens_out":4726,"duration_ms":42679,"significance":"If the causal claim held, this would be a significant demonstration that a large language model captures a subtle form-function interaction, with implications for both LM evaluation and theories of island constraints. The paper's strengths include zero-shot probing with temperature zero, ten repetitions per stimulus, replication on fresh stimuli to address contamination, public data, and a direct attempt at causal manipulation. However, the causal conclusion is currently undermined by confounds in Study 2 and by selective item analysis, so the central contribution remains the correlational finding. The correlational studies are informative and, with proper reporting of human correlations and model details, could constitute a solid contribution, but the paper's strongest claim ('confirms a causal relationship') is not yet supported.","major_comments":[{"comment":"The Emphasis condition differs from the No-emphasis condition not only in ALL-CAPS emphasis but also in the explicit instruction 'Please focus on the part in ALL CAPS.' This instruction is a demand characteristic that could raise ratings for the queried constituent independently of any representation of information structure. The reported causal claim therefore is not supported. The authors should add a control condition that includes an analogous instruction (e.g., 'Please focus on the part in lowercase' or 'Please focus on the whole sentence') or orthogonally vary emphasis and instruction in a 2x2 design.","section":"Study 2, Table 4"},{"comment":"The analysis is restricted to the 72 items with non-emphasis mean acceptability below 6. This is a post hoc selection on the control-condition outcome, and the reported effect (ß = 0.44, t = 8.48) may partly reflect regression to the mean. Please analyze the full set of items as the primary analysis, justify any exclusion criterion beforehand, and report the effect size for the full set; if the effect only appears in the selected subset, state that clearly.","section":"Study 2, Results"},{"comment":"The manuscript claims 'strong correlations with human judgments on the same stimuli' but no correlation coefficient, test statistic, or confidence interval is reported anywhere. Because the paper's framing compares GPT-4 with human BCI results, this correlation is essential. Please report the correlation between GPT-4 and human mean ratings (per item) for both backgroundedness and acceptability, and clarify whether the interaction slopes are quantitatively similar.","section":"Study 1a, Results/Introduction"},{"comment":"The exact model version is not specified (only 'GPT-4 API, Jan 2025'), and the random-effect structure is described only as 'the maximal random effect structure convergence allowed.' This hampers reproducibility. Please report the exact model identifier (e.g., gpt-4-0613 or gpt-4-turbo-2025-01) and the full model formula, including the random effects used in the ordinal model.","section":"Methods (proprietary model)"},{"comment":"The negation task is assumed to measure backgroundedness in GPT-4 just as it does in humans. Given that GPT-4's 'probably yes' answers might reflect shallow lexical or logical inferences (e.g., that a negated event still mentions the content), the paper should provide convergent validity evidence for this measure, for example by showing that the same backgroundedness judgments correlate with an independent information-structure probe or with the emphasis manipulation in Study 2 once the confound is removed.","section":"Study 1, Backgroundedness task"}],"minor_comments":[{"comment":"The phrase 'Here was ask' should be 'Here we ask.'","section":"Introduction"},{"comment":"The scale label '4 means probably yet' should be '4 means probably yes.'","section":"Table 2"},{"comment":"The text says 'three types of long-distance dependency constructions' and lists wh-Qs, discourse-linked questions, and relative clauses, but the Methods of Study 1a state that four types (including it-clefts) were collected; please reconcile the count and clarify which constructions were included in each study.","section":"General Discussion"},{"comment":"The reference 'Authors (2023)' in the description of the human study should be replaced with the proper citation 'Cuneo and Goldberg (2023).'","section":"Study 1a"},{"comment":"Figure 1 is described in the text but not visible in the manuscript; please ensure the figure is included and legible, as it is the only visual comparison of human and GPT-4 data.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about Study 2 is well founded. The confound between emphasis and the explicit instruction, combined with the post hoc item selection, means the causal claim is not supported as currently reported. The correlational studies are useful and likely worth publishing after the causal claim is either fixed with a cleaner experiment or substantially tempered. I would encourage the editor to ask for a revision that addresses the five major comments, especially the Study 2 issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: it is the first to show that GPT-4's explicit backgroundedness judgments predict its own acceptability ratings on long-distance dependencies, and it does so with two correlational studies and a new construction (it-clefts). That is a real result, and it makes the paper worth a serious referee. But the paper's causal claim, from Study 2, has a confound that the write-up doesn't resolve, and the analysis there has a selection issue that could manufacture part of the effect.\n\nWhat's new: the BCI interaction was already established for humans (Cuneo & Goldberg 2023; Lu et al. 2024). This paper shows GPT-4 replicates it, with zero-shot prompts, temperature 0, repeated queries, and open data. The interaction is significant in both 1a and 1b, and the design is genuinely careful about contamination. That's solid evidence that GPT-4's metalinguistic judgments track a form-function relationship, not just surface statistics. The authors also deserve credit for using the same negation task as the human studies, which makes the comparison meaningful.\n\nThe soft spots are mostly in Study 2. The Emphasis condition adds an explicit instruction to \"focus on the part in ALL CAPS,\" while the No-emphasis condition has no analogous instruction. That's a demand characteristic: telling the model to focus on a constituent can raise ratings on a following wh-question about it, independent of any representation of information structure. Second, the analysis is restricted to 72 items with non-emphasis mean below 6. That is a post hoc selection on the control condition outcome, and it can create or inflate the emphasis effect via regression to the mean. The paper acknowledges the selection but does not justify it. Those two issues together mean the causal claim, and the 'emergent understanding' language in the General Discussion, go beyond what Study 2 actually shows.\n\nThe correlational studies are not tainted by these problems, so the central finding stands. But the causal story needs a redesign: ideally the same instruction in both conditions, or no explicit instruction, and an analysis on the full item set or with a pre-registered criterion. The missing model version and random-effect specification are minor but should be reported.\n\nWho is this for? Psycholinguists interested in island effects, and anyone evaluating LLMs with metalinguistic prompts. I would cite the correlational results. It deserves peer review, with the expectation of major revision.","headline":"Solid correlational evidence that GPT-4 tracks information-structure-to-acceptability mappings; the causal Study 2 is confounded by an explicit instruction and post-hoc item selection.","tokens_in":9056,"tokens_out":2799,"would_cite":true,"duration_ms":26795,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4 treats backgrounded content as islands, exactly as humans do.","keywords":["long-distance dependencies","island constraints","information structure","backgroundedness","GPT-4","acceptability judgments","metalinguistic judgment","constructions"],"falsifier":"Compare GPT-4's backgroundedness ratings on items where human presupposition judgments diverge from logical entailment (for example, negated clausal complements whose content is mentioned but not presupposed); if backgroundedness no longer predicts LDD acceptability on those items, the claimed role of information structure is not supported. Replacing the negation task with a focus-based measure, such as question-answer congruence, and checking whether the LDD predictability survives would provide a second decisive test.","tokens_in":8107,"feed_emoji":"🧠","tokens_out":8607,"duration_ms":76756,"temperature":0.7,"pith_summary":"This paper claims that GPT-4, with no in-context examples and no fine-tuning, reproduces a subtle relationship recently established in English speakers: the more backgrounded a constituent is in a canonical sentence, the less acceptable it is in a corresponding long-distance dependency such as a wh-question or relative clause. Using the same negation-based backgroundedness task and the same acceptability-rating prompts used with humans, the paper shows that GPT-4's backgroundedness judgments predict its own acceptability ratings on four types of LDDs, and that this prediction is stronger for LDDs than for the base sentences themselves. A second study makes the relationship causal: when a preceding context emphasizes the to-be-queried constituent, GPT-4 rates the following wh-question as more acceptable. The result matters because it suggests that large language models are not limited to shallow syntactic mimicry or memorized exemplars, but represent the same gradient, discourse-driven constraint on extraction—'backgrounded constituents are islands'—that has been proposed for human grammars.","feed_headline":"GPT-4 mirrors humans: backgrounded phrases resist extraction","feed_subtitle":"Zero-shot ratings reproduce the information-structure effect behind island constraints.","key_machinery":"The load-bearing mechanism is the BCI principle—'Backgrounded Constituents are Islands'—which states that a constituent's unavailability for long-distance dependencies scales with how backgrounded its content is. The paper operationalizes backgroundedness with the negation task: after a negated sentence, a question about the key constituent is answered 'probably yes' if that content survives negation, meaning it is presupposed or backgrounded rather than at-issue. This gradient backgroundedness score, collected on canonical base sentences, is then correlated with independently elicited acceptability ratings on LDDs built from those same sentences. In Study 2 the machinery becomes causal: a context sentence with lexical emphasis on the to-be-queried constituent increases that constituent's prominence and, in turn, GPT-4's acceptability rating of the subsequent wh-question.","core_discovery":"On the paper's own terms, the central discovery is that GPT-4's explicit metalinguistic judgments about information structure predict its independent acceptability ratings on long-distance dependency constructions, in the direction predicted by the Backgrounded Constituents are Islands (BCI) principle: constituents rated more backgrounded in a base sentence receive lower acceptability ratings when they are extracted in LDDs. This holds for wh-questions, discourse-linked questions, relative clauses, and it-clefts, and it holds with newly generated stimuli designed to rule out training-data contamination. Study 2 goes further: when a preceding context sentence emphasizes the to-be-queried constituent, GPT-4 gives higher acceptability ratings to the following wh-question than when the same context is presented without emphasis. The paper concludes that GPT-4 captures a systematic, causal relationship between information structure and syntactic acceptability that parallels human psycholinguistic results.","pith_inferences":["A testable extension: if the BCI interaction is tied to scale, smaller or open-weight models should show a weaker or absent effect; a graded pattern across model sizes would clarify whether this is an emergent property of language-model training objectives.","The same prominence manipulation could be probed in generation: GPT-4 may be more likely to produce wh-questions that extract constituents made prominent in the preceding context, offering a behavioral test beyond explicit ratings.","The negation-task assumption can be stress-tested by comparing GPT-4's ratings with formal presupposition-projection diagnostics; divergence would mark the boundary of the model's information-structure sensitivity."],"forward_implications":["If the pattern holds, GPT-4 can serve as a zero-shot informant for gradient acceptability judgments, letting researchers probe island constraints and information structure without large human norming samples.","The BCI principle gains a new form of evidence: a model with no explicit linguistic rules reproduces the interaction, strengthening discourse-functional accounts over purely syntactic ones.","Because Study 1b used fresh stimuli and added it-clefts, the effect transfers to a construction not present in the original human data, arguing against contamination by memorized examples.","The causal emphasis effect from Study 2 implies that acceptability judgments in GPT-4 are context-sensitive, not static ratings of sentence form alone."],"supporting_citations":[{"why":"Supplies the 144 base sentences and the human acceptability data that Study 1a replicates with GPT-4.","marker":"Cuneo & Goldberg, 2023"},{"why":"Provides the negation-task measure of backgroundedness and the evidence that it predicts adjunct island status.","marker":"Namboodiripad et al., 2022"},{"why":"States the Backgrounded Constituents are Islands principle that the paper tests in GPT-4.","marker":"Goldberg, 2006, 2013"},{"why":"Contributes the emphasis-manipulation design that Study 2 adapts to establish a causal effect.","marker":"Lu, Pan, and Degen, 2024"},{"why":"Shows the same causal emphasis effect for single-conjunct wh-questions, supporting the manipulation's generality.","marker":"Fergus et al., 2025"}],"fun_headline_variants":["GPT-4 mirrors humans: backgrounded phrases resist extraction","Backgrounding predicts GPT-4's island constraints","Context emphasis raises GPT-4's extraction acceptability","Zero-shot GPT-4 matches human island judgments","Info structure governs GPT-4's syntax acceptability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that GPT-4's 'probably yes' answers on the negation task reflect genuine presupposition or backgroundedness, rather than a shallow inference that a negated sentence still mentions the content; if the latter were true, the correlation with LDD acceptability could be driven by an unmeasured third factor.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 mirrors humans: backgrounded phrases resist extraction","Backgrounding predicts GPT-4's island constraints","Context emphasis raises GPT-4's extraction acceptability","Zero-shot GPT-4 matches human island judgments","Info structure governs GPT-4's syntax acceptability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":2005,"prompt_tokens":934,"completion_tokens":1071,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":998}},"tokens_in":550,"tokens_out":1071,"duration_ms":10341,"temperature":1.0,"reasoning_tokens":998,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:42:04.674385+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare GPT-4's backgroundedness ratings on items where human presupposition judgments diverge from logical entailment (for example, negated clausal complements whose content is mentioned but not presupposed); if backgroundedness no longer predicts LDD acceptability on those items, the claimed role of information structure is not supported. Replacing the negation task with a focus-based measure, such as question-answer congruence, and checking whether the LDD predictability survives would provide a second decisive test.","supporting_citations":[],"review_version":1}