{"id":"bef9e043-7acf-471b-ad3b-a04ecc6ab705","arxiv_id":"2506.04534","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs separate broad senses of \"just\" such as temporal and adjective, but on naturalistic sentences they struggle to distinguish the fine-grained discourse-particle senses.","lead":"This paper tests whether large language models can tell apart the different meanings of the small word \"just\", such as \"only\", \"recently\", or \"emphatic\", using sentences labeled by expert linguists. It finds that models succeed on broad distinctions but miss subtle ones, a result that matters for anyone deploying AI systems that need to understand conversational nuance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subtitle ground-truth labels are the load-bearing premise; without reported inter-annotator agreement, near-chance accuracy may reflect labeling noise rather than a semantic deficit.","rationale":"The reader identified the validity of the ground-truth sense labels as the most fragile premise, and my reading agrees: the headline claim is a measured mismatch between model predictions and expert labels, so any systematic error in those labels directly contaminates both the sense-labeling accuracies and the pairwise separation effect sizes. Other concerns, such as the use of t-tests on non-independent sentence pairs and the absence of confidence intervals, are real but are fixable statistical reporting issues; they affect how strongly the data support the claim, not whether the data measure what the paper says they measure. The label-reliability concern is more load-bearing because it threatens construct validity: if many subtitle items are genuinely ambiguous or the taxonomy over-splits senses, then the near-chance performance on naturalistic sentences would be an artifact of the evaluation instrument, not a semantic deficit in LLMs. The paper's own acknowledgment of ambiguity in examples like \"I just saw Nancy\" makes this risk concrete. The hand-constructed data and the bat/bank control provide independent support for the method and for coarse sense separation, which is why I would not reject or mark the paper unverdictable; however, the central negative claim should be accepted only conditional on demonstrating label reliability. Since the reader's verdict is already CONDITIONAL and this concern is the same one, I recommend no change to the verdict.","tokens_in":11355,"tokens_out":3383,"duration_ms":43157,"concrete_test":"Release item-level annotation counts for all 149 subtitle sentences and compute Fleiss' kappa or Krippendorff's alpha over the annotator responses, restricting to items with at least three annotations; then have two or three independent semanticists blind-label the 90 hand-constructed sentences. If agreement is high (alpha > 0.8) and the hand-constructed labels replicate, the concern is resolved. If agreement is moderate or low, re-run the accuracy and pairwise separation analyses using only high-agreement items and compare model scores against the labelable ceiling implied by annotator agreement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central negative result hinges on treating expert and annotator labels as ground truth. Section 3 asserts that \"all sentences in both datasets have a strong primary reading\" either by construction or by annotator agreement, but no agreement statistic is reported. The annotation procedure is also underspecified: a variable subset of eight annotators labeled each of the 149 subtitle sentences, and when there was disagreement, two senior annotators' labels were \"both considered regardless of agreement.\" This leaves unclear how a final gold label was chosen and whether the claimed consensus is strong. With a skewed label distribution (60 Exclusive, 33 Temporal, 22 Unexplanatory, 21 Emphatic, 12 Unelaboratory, 1 Adjective) and only 149 items, a modest fraction of mislabeled items would materially depress measured accuracy: if 20% of gold labels are wrong, the achievable ceiling is roughly 0.8, and a model scoring 0.45 could be close to the labelable ceiling rather than deficient. The hand-constructed data mitigate this for the pairwise experiment because 15 sentences per sense were designed to be unambiguous by a native-expert linguist, but those labels are also single-annotator and unchecked. The pairwise separation statistics are likewise computed against these labels, so if the taxonomy over-splits subtle discourse-particle senses, effect sizes would be attenuated. Thus the claim that even the largest models fail to capture just's richness is not airtight until label reliability is quantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether instruction-tuned LLMs correctly distinguish the fine-grained senses of the English discourse particle \"just\" (exclusive, unelaboratory, unexplanatory, emphatic, plus temporal and adjective controls). Using 90 hand-constructed sentences and 149 OpenSubtitles sentences annotated by expert semanticists, the authors run two experiments: (1) prompted sense labeling with definitions and examples, and (2) pairwise judgments of whether two occurrences of \"just\" are used the same way. They report that most models exceed uniform chance on the hand-constructed data but perform near the majority-class baseline on naturalistic subtitles, that context does not help, and that pairwise separation is clear only for coarser senses such as adjective and temporal. The authors conclude that LLMs have basic but incomplete knowledge of discourse-particle senses.","tokens_in":11560,"tokens_out":3527,"duration_ms":44456,"significance":"If the results hold, this is a valuable behavioral case study with expert-constructed and expert-annotated data, two complementary experimental paradigms, and a cross-scale model comparison that includes both open and commercial-style instruction-tuned models. The explicit release of code and data is a strength, as are the bat/bank control experiments that validate the pairwise methodology. The paper extends prior work on discourse relations to a class of understudied function words, and its negative result is falsifiable. However, the central negative claim is load-bearing on the reliability of the gold labels and on the statistical treatment of small, imbalanced samples, both of which need strengthening before the conclusions can be considered airtight.","major_comments":[{"comment":"The subtitle-label ground truth is the load-bearing premise for the central negative result, but no inter-annotator agreement statistic is reported and the adjudication procedure is underspecified: the text states that a variable subset of eight annotators labeled each sentence and that when there was disagreement, \"two additional senior annotators, whose labels were both considered regardless of agreement\" were used, but it never explains how a single gold label was derived from this process. With only 149 items and a skewed distribution (60 Exclusive, 33 Temporal, 22 Unexplanatory, 21 Emphatic, 12 Unelaboratory, 1 Adjective), even a modest fraction of mislabeled items materially lowers the achievable accuracy ceiling: if 20% of gold labels are wrong, the ceiling is roughly 0.8, and a model scoring 0.45 could be near that ceiling. The authors should report agreement statistics (e.g., Fleiss' kappa), per-item confidence ratings, and a clear gold-label adjudication rule, and they should make the annotator-by-item label matrix available so that the ceiling is estimable.","section":"Section 3 and §4.2"},{"comment":"The claim that all models except Llama-3.2-1B show \"significant separation\" (p < .005) rests on Welch t-tests computed on 1260 same-sense pairs and 6750 different-sense pairs, but these pairs are not independent: they share common sentences, so the effective sample size is far smaller than the pairwise counts and the p-values are anti-conservative. The same issue affects the reported Cohen's d effect sizes, which are not accompanied by confidence intervals. The authors should either use a cluster-robust test or a mixed-effects model with sentence- or pair-level random effects, or a permutation test that preserves the dependence structure, and they should report confidence intervals for the effect sizes. The qualitative heatmap patterns are suggestive, but the quantitative significance claim needs a sounder statistical basis.","section":"Section 5.2, Table 5"},{"comment":"Accuracy is reported as point estimates without confidence intervals or explicit tests against the chance baselines. For the 149-item subtitle set, the difference between a model at 0.45 and the majority-class baseline of 0.403 may not be meaningful, and the claim that models are \"at chance\" for several conditions should be supported by binomial or bootstrap confidence intervals per model and per condition. Without such intervals, it is impossible to judge which model-by-condition differences are reliable, and the pattern of \"near chance\" on subtitles versus above-chance on hand-constructed data is not quantified in a way that supports the paper's central conclusion.","section":"Section 4.2, Figure 2"},{"comment":"The hand-constructed sentences are described as having been \"carefully created by an expert\" who is a graduate linguist and native speaker, but no second annotation, adjudication, or agreement measure is reported for these 90 items. Since the pairwise experiment in §5.2 relies entirely on these labels as ground truth, any ambiguous or mislabeled items would attenuate the measured same-sense versus different-sense separation, particularly for the fine-grained target senses that are the focus of the paper. A validation pass by a second annotator, or at least a public release of the items with confidence ratings, would strengthen the interpretation of the effect sizes in Table 5.","section":"Section 3 and §5.2"}],"minor_comments":[{"comment":"The sense-labeling prompt uses the label \"Exclusionary\" for what the rest of the paper calls \"Exclusive\"; the authors should harmonize this terminology to avoid confusion in reproducing the exact prompt.","section":"Appendix F, Figure 6"},{"comment":"The sentence \"strong speaker consensus on the reading of an occurrence of just does remove more ambiguous sentences from out data\" contains a typo: \"out data\" should be \"our data.\"","section":"Section 3"},{"comment":"The notation is inconsistent across model names (e.g., \"Llama3.2 1B\" in Table 5 versus \"Llama-3.2-1b\" in the text and figures); the authors should standardize model identifiers.","section":"Section 5.2 and Table 5"},{"comment":"The \"ideal\" heatmap is embedded in the same row as the model heatmaps, which makes it easy to mistake it for a model output; placing it separately or adding a clear panel label would improve readability.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the journal and the core empirical pattern is plausible and interesting. The main concern is that the subtitle-label reliability and the non-independence of the pairwise statistics are load-bearing for the strongest claims; these are fixable with additional analysis and reporting rather than requiring new experiments. I would also gently note that because one of the authors is a proponent of the unified account cited for the taxonomy, the annotation guidelines should make clear whether annotators were forced to map the theory's categories onto each occurrence; this does not constitute circularity, but it affects how the negative result should be interpreted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe bottom line: this is an honest, well-scoped study of whether LLMs internalize the fine-grained senses of the English particle 'just'. The new contribution is real—a small expert-labeled dataset, a log-probability sense-labeling protocol, and a pairwise comparison with bat/bank controls. The two experiments complement each other: the labeling task tests metalinguistic knowledge, the pairwise task tests implicit similarity. The result that models separate adjective and temporal senses but struggle with the four discourse-particle senses is consistent across both methods and model families. That is worth knowing.\n\nThe paper does several things right. The controls are sensible—bat and bank show clean separation, so the pairwise method is not vacuous. Using log-probabilities rather than parsing generations avoids a common fragility. Code and data are available, and the limitations section acknowledges the metalinguistic prompting concern. I appreciate that the authors picked a well-studied word from formal semantics and built evaluation around its attested senses.\n\nThe soft spot is the load-bearing ground truth. The subtitle annotations have no reported inter-annotator agreement, the procedure for resolving disagreement is under-specified ('both labels considered regardless of agreement' is not a recipe for a gold label), and the final sense distribution is skewed. With 149 items, a modest fraction of noisy labels could push the near-chance accuracy down substantially. The hand-constructed data mitigates this for the pairwise experiment, but those labels come from a single annotator. The paper itself notes 'I just saw Nancy' is ambiguous, so the 'strong primary reading' claim needs evidence. The statistical testing also has a known issue: the Welch t-tests on same-sense vs different-sense pairs treat overlapping pairs as independent. The effect sizes are still informative, but the p-values are optimistic.\n\nNone of this kills the main finding—models clearly don't fully capture the fine senses—but the magnitude of the deficit is uncertain. Adding confidence intervals, reporting IAA, and either correcting the statistics or softening the pairwise significance claims would make the argument much stronger.\n\nWho is this for? People working on discourse semantics in NLP, and anyone building evaluations for polysemous function words. It deserves a serious referee; the topic is important, the design is thoughtful, and the dataset is a reusable resource. I'd recommend sending it to review rather than rejecting, with requests for the label reliability analysis and statistical fixes.","headline":"A solid, honest study of LLM sensitivity to the senses of 'just' whose main conclusion is plausible but currently over-claimed because the gold labels are not shown to be reliable.","tokens_in":12142,"tokens_out":3095,"would_cite":true,"duration_ms":35308,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that instruction-tuned LLMs, although able to separate broad uses, fail to reliably distinguish the fine-grained discourse-particle senses of the English word 'just' in naturalistic text.","keywords":["discourse particles","just (English)","LLM sense disambiguation","formal semantics","pragmatics","polyfunctionality","metalinguistic prompting","OpenSubtitles"],"falsifier":"Have independent semanticists label the 90 hand-constructed and 149 subtitle sentences and calculate inter-annotator agreement; then score models against each annotator's label rather than the majority. If model accuracy equals the human annotation ceiling on ambiguous items, or if the expert labels themselves diverge on a large share of sentences, then the near-chance performance is a labeling artifact rather than a semantic deficit in the models.","tokens_in":11125,"feed_emoji":"💬","tokens_out":9664,"duration_ms":86898,"temperature":0.7,"pith_summary":"The paper asks whether instruction-tuned large language models know the fine-grained senses of the English discourse particle 'just'—the exclusive, unelaboratory, unexplanatory, emphatic, temporal, and adjective uses that formal semantics distinguishes. Using expert-constructed and expert-annotated sentences, it shows that models can tell apart the broad categories (adjective and temporal) but do not reliably identify the four nuanced discourse-particle senses, even when sense definitions and examples are provided in the prompt. On naturalistic movie-subtitle sentences, accuracy falls close to chance for most models, and adding conversational context does not help. A pairwise judgment task confirms the pattern: models separate clearly distinct senses like the controls 'bat' and 'bank', but only weakly separate the target senses. If right, this means current LLMs lack an internalized, fine-grained semantics for a frequent and discourse-critical word.","feed_headline":"LLMs miss fine-grained senses of 'just'","feed_subtitle":"With expert definitions and context, models stay near chance on real sentences; larger models don't close the gap.","key_machinery":"Two probing tasks carry the argument. The first is a metalinguistic sense-labeling probe: the model assigns one of six sense labels by taking the label with the highest conditional probability after a prompt that defines each sense with examples. The second is a pairwise same-use probe: for each pair of sentences the model rates whether the two occurrences of 'just' are used the same way, and the normalized log-probability difference between 'Yes' and 'No' forms a heatmap whose block structure measures sense separation. The load-bearing resources are the two expert-labeled datasets—90 hand-constructed unambiguous sentences and 149 OpenSubtitles sentences—and the control items ('bat', 'bank', adjective and temporal senses) that show the probes can detect clear distinctions when they exist.","core_discovery":"The central discovery is a measured gap between the coarse and fine senses of 'just' in LLMs. In the labeling task, models are scored by the conditional probability of each sense label; all models except the 1B one beat chance on 90 hand-built unambiguous sentences, but accuracy drops by about 0.24 on 149 subtitle sentences and does not recover when two previous utterances are added as context. In the pairwise task, the model compares sentences and gives the difference in log-probability of 'Yes' and 'No'; same-sense pairs are rated higher than different-sense pairs for every model except the smallest, but the effect sizes for the four target senses are small, while adjective, temporal, and the control words 'bat' and 'bank' separate cleanly. The authors read this as evidence that LLMs have a coarse, partially correct grasp of 'just' but have not internalized the discourse-particle distinctions that formal semantics identifies.","pith_inferences":["The result likely extends beyond 'just' to other polyfunctional discourse particles such as 'actually', 'even', 'like', and 'though'; a direct test would be to rerun the same two probes on a small inventory of such words.","Because the paper notes that intonation often disambiguates 'just' in speech, text-only models may be operating with genuinely less information; an extension would be to test speech or prosody-aware models on the same sentences.","The absence of a context benefit may reflect the instruction-following setup rather than the models' internal representations; a direct sentence-surprisal measure could reveal sensitivity that metalinguistic prompting misses, a caveat the paper itself flags.","A human-model alignment analysis—scoring models by agreement with individual annotators instead of majority labels—would separate genuine semantic deficits from ambiguity in the gold standard."],"forward_implications":["Scaling model size alone will not close the gap: 70B models perform near chance on naturalistic sentences and show only weak separation of the four target discourse senses.","Providing conversational context is not a cure; model accuracy did not improve and often fell when two prior utterances were added.","The pairwise same-use probe is sensitive enough to detect clearly separated senses (adjective, temporal, and the control words bat and bank), so the weak signals for 'just' point to a property of the particle's semantics rather than a broken measurement.","Systems that rely on LLMs for discourse interpretation—dialogue, summarization, reasoning about speaker intent—should expect errors where a polyfunctional particle like 'just' carries the meaning.","Because the task uses labels from formal semantic theory, the result also constrains what 'understanding' means for function words: coarse category knowledge can coexist with missing fine-grained internalized senses."],"supporting_citations":[{"why":"Provides the six-way sense taxonomy and the example sentences used in the hand-constructed dataset and prompt definitions.","marker":"Warstadt (2020)"},{"why":"Supplies the OpenSubtitles2018 corpus from which the 149 naturalistic sentences were drawn and annotated.","marker":"Lison et al. (2018)"},{"why":"Gives the unified account of 'just' with twelve senses that motivates the taxonomy the paper tests.","marker":"Deo and Thomas (2025)"},{"why":"Defines the exclusive particle semantics that grounds the exclusive sense of 'just'.","marker":"Coppock and Beaver (2014)"},{"why":"Provides the minicons toolkit used to compute the conditional log-probabilities for label scoring.","marker":"Misra (2022)"},{"why":"Prior evidence that LLMs struggle with discourse relations, setting the expectation the paper extends to discourse particles.","marker":"Chan et al. (2024)"},{"why":"The paper's limitation discussion cites it to flag that metalinguistic prompting may underestimate LLM abilities.","marker":"Hu and Levy (2023)"}],"fun_headline_variants":["LLMs stumble on the many meanings of 'just'","The many senses of 'just' confound LLMs","LLMs can't pin down the fine senses of 'just'","Why 'just' is just hard for LLMs","LLMs fail to grasp the nuanced 'just'"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that every sentence in the two datasets has one correct intended sense of 'just'; if many sentences are genuinely ambiguous or the six-way taxonomy splits natural uses too finely, then low model accuracy could reflect disagreement in the labels rather than a real gap in the models' understanding.","fun_headline_variants_meta":{"raw":{"variants":["LLMs stumble on the many meanings of 'just'","The many senses of 'just' confound LLMs","LLMs can't pin down the fine senses of 'just'","Why 'just' is just hard for LLMs","LLMs fail to grasp the nuanced 'just'"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000532,"raw_usage":{"total_tokens":2508,"prompt_tokens":837,"completion_tokens":1671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":1590}},"tokens_in":453,"tokens_out":1671,"duration_ms":12459,"temperature":1.0,"reasoning_tokens":1590,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:39:48.148478+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent semanticists label the 90 hand-constructed and 149 subtitle sentences and calculate inter-annotator agreement; then score models against each annotator's label rather than the majority. If model accuracy equals the human annotation ceiling on ambiguous items, or if the expert labels themselves diverge on a large share of sentences, then the near-chance performance is a labeling artifact rather than a semantic deficit in the models.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the OpenSubtitles2018 corpus from which the 149 naturalistic sentences were drawn and annotated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the unified account of 'just' with twelve senses that motivates the taxonomy the paper tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the exclusive particle semantics that grounds the exclusive sense of 'just'."}],"review_version":1}