{"id":"1d6045d7-2ec9-4b14-90c8-2a3b812af06e","arxiv_id":"1908.10747","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A methodology paper that formalizes tasks, worlds, and games in NLP and argues that progress claims require explicit assumptions about decomposable language capabilities.","lead":"This paper proposes formal definitions for 'language task', 'micro-world', and 'language game' to clarify how NLP research claims progress. It argues that without specifying how a task relates to language capabilities, new datasets and benchmarks do not automatically count as progress toward general language competence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the paper's central normative claim does not depend on a fully specified set of language capabilities.","rationale":"The reader's verdict ACCEPT with HIGH confidence is justified. The paper does what it claims: it defines task, world, game, and provides a structured argument that task-based NLP research needs explicit claims about language capabilities. The strongest claim is normative and follows from the definitions: progress toward modelling general language competence must be argued through a relation between the task and that goal. The paper acknowledges the main fragility (Section 3.1), and the remaining informality of the 'involves' relation is a call for future work, not a flaw. There are no empirical claims to check, no code, and no formal proofs, but for a position piece the argument is internally coherent. I considered whether the unformalized C_L makes the recommendation unactionable, but the paper's own examples (VQA, NLI) show how one at least starts: by claiming specific capabilities such as compositionality, grounding, or inference. The recommendation to make such claims explicit and contestable is a meaningful methodological contribution even if the full decomposition of C_L is unknown. Hence no change to the verdict.","tokens_in":10971,"tokens_out":7368,"duration_ms":79876,"concrete_test":"Formalize the 'involves' relation between a task T and a capability c without assuming a partition of C_L into per-task subsets; then check whether the Section 5 conclusion 'motivating those in themselves... requires making claims about language capabilities' still follows from the Section 3.1 capability argument. If the conclusion follows without the partition, the reader's weakest assumption is not load-bearing; if it does not, the paper's central claim needs the decomposability assumption after all.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption concerns the existence and decomposability of a well-defined set C_L of competent-language-user capabilities, and the paper itself flags this in Section 3.1 ('it seems difficult to properly motivate a task without relying, at least implicitly, on assumptions about how C_L decomposes'). I do not think this is a load-bearing objection to the central claim. The paper's argument is conditional: if a task is presented as progress toward modelling general language competence, then its motivation must make explicit which capabilities it is claimed to involve and why that claim is contestable. This demand remains coherent even if C_L is only partially known or not cleanly decomposable; the request is for explicit, falsifiable assumptions, not for a finished theory of language competence. The paper explicitly directs researchers to linguistics and cognitive psychology for the needed constructs, which is a reasonable (if non-trivial) program. The one genuine gap is that the 'involves' relation between tasks and capabilities is left informal, so the recommendation could in principle be satisfied by vacuous claims; however, the paper already tries to rank specificity from 'capability to do T' to stronger falsifiable claims, and offers this as a methodological starting point rather than a completed framework. This is a limitation, not an internal inconsistency.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that current NLP research habitually introduces new tasks, datasets, worlds, and games while leaving implicit what such contributions are progress toward. The author defines a language task as a mapping between input and output/action spaces involving natural language, a micro-world/environment as a mapping from actions to environmental responses, and an interaction game as a setting with repeated, connected language tasks regulated by turn-taking and evaluation rules. The central normative claim is that motivating a task or game requires making explicit claims about how the capabilities C_T it involves relate to the larger set C_L of capabilities of a competent language user; without such claims, a new task or dataset does not by itself constitute progress. The paper distinguishes three modes of progress (better models, better datasets, better tasks/games), analyzes their different argumentative burdens, and recommends a stronger connection to linguistics and cognitive psychology to make capability claims contestable. Formal definitions are provided in an appendix.","tokens_in":11164,"tokens_out":3401,"duration_ms":38844,"significance":"If the argument is accepted, it gives researchers a vocabulary for evaluating task and dataset contributions beyond benchmark scores, and it sharpens a critique that is often voiced only informally. The paper's main strength is that it makes an implicit argument structure explicit: progress claims rest on assumptions about the decomposition of language competence, and those assumptions should be stated so they can be challenged. The treatment of the VQA language-bias example and the discussion of dataset validity are concrete and instructive. The paper also honestly flags its own limitations, most notably in Section 3.1, where it acknowledges that motivating a task requires assumptions about how C_L decomposes. The paper contains no technical derivations or experiments, so its contribution is conceptual and methodological; it is a well-targeted piece of scientific criticism rather than a new empirical result. The central claim is conditional and remains coherent even without a fully specified theory of language capabilities.","major_comments":[],"minor_comments":[{"comment":"The ranking of capability claims from the trivial \"task T involves the capability to do task T\" to stronger falsifiable claims is intuitive, but the \"involves\" relation itself is never defined formally; a short worked example showing how a specific task would be paired with a non-trivial, falsifiable capability claim would make the central recommendation easier to apply.","section":"Section 3.1"},{"comment":"There are several typographical errors: \"priviledged\" in Section 2.3 should be \"privileged\", \"Targetting\" in Section 4.2 should be \"Targeting\", and \"diagramm\" in Section 4.3 should be \"diagram\".","section":"Section 2.3 and Section 4.2"},{"comment":"The condition \"with either the states in S or the actions in A (or both) having as part natural language expressions\" is grammatically awkward and could be misread as requiring states or actions themselves to contain language expressions; a cleaner formulation would describe a function from states or actions to natural language expressions.","section":"Appendix A.1, Definition 1"},{"comment":"The statement that tasks \"are only grounded (to their left in the diagram) by capabilities\" relies on the spatial layout of Figure 2; consider stating the direction of grounding textually without reference to the diagram's left side, since the figure may be rendered or read differently.","section":"Section 4.3"}],"recommendation":"minor_revision","confidential_remarks":"This is a methodology/position paper rather than a technical contribution, and it fits the scope of a venue that publishes critical and conceptual work in NLP. The central argument is sound, and the manuscript's limitations are largely self-acknowledged. The minor comments above concern presentation and clarifications that can be addressed without changing the substance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper deserves a serious read even though it contains no new empirical results. What it does instead is give a clean formal account of what a language task is, distinguishes tasks from micro-worlds and interaction games, and then argues that any new task or dataset only counts as progress if its relation to language capabilities is made explicit. The appendix definitions—task as a tuple, environment as a mapping, interaction game with a disinterested Nature player—are genuinely useful and carry over into the main text. That alone separates it from the usual 'benchmarks are bad' commentary.\n\nThe strongest part is the decomposition of progress into better models, better datasets, and better tasks, and the observation that only the first has a straightforward evaluation. The VQA bias example is a good, well-chosen illustration of a dataset being invalid relative to its task description. The paper is also honest about its own load-bearing assumption: it flags in Section 3.1 that motivating a task requires assumptions about how the set of capabilities C_L decomposes. That self-awareness helps credibility.\n\nSoft spots are proportionate. The 'involves' relation between a task and a capability is left informal, so the main recommendation could in principle be satisfied by a vacuous claim like 'this task requires the capability to do this task.' The paper acknowledges this and tries to rank specificity, but it remains a genuine gap. Also, the claims about games being less well-exemplified by datasets are supported by anecdote rather than systematic evidence. Neither of these sinks the argument. The central normative claim—that task introductions should state which capabilities they involve and why—stands even if C_L is only partially known. The stress-test note is right: the argument is conditional and does not require a finished theory of language competence.\n\nThis paper is for anyone designing or reviewing NLP tasks, and for the field's ongoing discussion about benchmark culture. It deserves a serious referee; the formalizations and the taxonomy of progress modes are a real contribution even if the empirical claims are modest. I would suggest the authors tighten the 'involves' relation and give more than one worked example of a non-vacuous capability claim, but neither is a blocker. Send it to peer review.","headline":"A genuinely useful methodological critique that gives NLP researchers a cleaner vocabulary for task design, and the formal appendix does real work.","tokens_in":11621,"tokens_out":1644,"would_cite":true,"duration_ms":18330,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new NLP task or dataset is only progress if it states which language capabilities it exercises; otherwise, the paper argues, the contribution is vacuous.","keywords":["language tasks","language games","NLP methodology","benchmark evaluation","language competence","dataset validation","capability decomposition","intelligent agent environments"],"falsifier":"Run a large battery of diverse language tasks with human participants and factor-analyze the performance matrix. If no stable task-specific capability components emerge — that is, a single general language factor explains nearly all variance — then the paper's assumption that the overall capability set CL decomposes into task-linked subsets CT loses its empirical grounding.","tokens_in":10798,"feed_emoji":"🧠","tokens_out":9040,"duration_ms":89979,"temperature":0.7,"pith_summary":"This paper examines the common NLP practice of contributing a new task and dataset and asks why such contributions are assumed to be progress. It proposes precise definitions of a language task, a micro-world or environment, and a language game, and it argues that a new task or dataset counts as progress toward modeling general language competence only if its relation to the capabilities of a competent language user is made explicit. The author's central point is that motivating a task in itself requires claims about which capability subset the task involves; how good such motivation is depends on how precisely and falsifiably those claims are stated. If the paper is right, benchmark introductions should come with explicit capability analyses, and dataset deficiencies can be detected by showing that a model solves the task without information the task description treats as essential.","feed_headline":"New NLP tasks only count as progress if they say which language ability they test","feed_subtitle":"Without that link, a new dataset or benchmark cannot claim to move us toward modeling general language competence.","key_machinery":"The machinery is a set of formal tuple definitions — language task $(S, A, L, D_T)$, micro-world/environment $(S, A, E, R, D_W)$, interaction game $(P, A, o, T, E, D_G)$ with a disinterested player Nature — together with the 'involves' relation between a task and a capability subset. The definitions give a language to compare tasks, datasets, environments, and games, and the 'involves' relation supplies the yardstick: a contribution is progressive only when the claimed capability subset is specified precisely enough to be tested and is related to the overall capability set of a competent language user.","core_discovery":"The central claim is that a new NLP task or dataset only constitutes progress if the relation between the task and the capabilities of a competent language user is stated explicitly; otherwise, the contribution is vacuous. The paper grounds this claim by defining a language task as a mapping from states to actions, at least one side involving natural language expressions, conforming to a task description; an environment as an action-to-state mapping conforming to a world description; and an interaction game as players, action spaces, an observability function, a turn-taking rule, an evaluation rule, and a game description. It then argues that any task motivation presupposes a set CL of capabilities of a competent language user and a subset CT that the task exercises. The strength of the motivation is ranked by how precisely CT is specified and by two dimensions: separability (whether CT can be handled independently of the rest of CL) and exhaustivity (whether the task exercises all of CT). A claim like 'the task involves the capability to compute syntactic structure' is strong because it could be wrong; a claim like 'the task involves the capability to do the task' is trivial.","pith_inferences":["A natural next step the author leaves implicit is a 'capability sheet' for each dataset or task, listing the claimed CT, its separability and exhaustivity, and the evidence that would falsify the claim; this would make the 'involves' relation contestable by inspection.","If the capability-decomposition view is correct, then transfer learning provides a direct test: a model trained on a task that genuinely engages CT should improve performance on another task engaging the same CT; failures of transfer would indicate the capability claim was wrong.","The framework could be applied to evaluate 'AI-complete' claims: saying a task is AI-complete amounts to claiming CT = CL, and by the paper's ranking that is the strongest and least likely kind of capability claim, not a casual rhetorical flourish.","Operationalizing 'closeness to unrestricted situated language interaction' could yield a graded benchmark-construction guide (e.g., increasing the number of interacting players, hiding observations, or lengthening action sequences), although the paper does not specify such a metric."],"forward_implications":["If the paper is right, a new task proposal should state, in falsifiable terms, which subset of competent-language-user capabilities the task engages, and how that subset relates to subsets of existing tasks.","A dataset can be rejected as an unsatisfactory exemplification of a task when a model solves it without information the task description deems crucial, as happened when visual question answering was solved without visual input.","Progress in tasks themselves means arguing that the new task's capability subset is more specific, more separable, or more exhaustive than the old one, not merely that a new dataset has been collected.","Language games can be assessed by how close they come to unrestricted situated language interaction, the natural upper bound for language-game design.","Probing results on trained models (for example, finding syntactic structure inside a network) should be read as evidence about whether a postulated capability is needed for the task, rather than as standalone architectural curiosities."],"supporting_citations":[{"why":"Introduced visual question answering, the paper's central case of a task motivated by broad capability claims whose dataset later proved solvable without visual input.","marker":"Antol et al. (2015)"},{"why":"Built a less biased VQA corpus, demonstrating the dataset-improvement mode of progress.","marker":"Goyal et al. (2017)"},{"why":"Introduced the SHAPES dataset to push compositionality, an example of task improvement aimed at a specific capability.","marker":"Andreas et al. (2016)"},{"why":"Historical notion of micro-worlds and the project of knitting them together; the paper uses this to frame environments and the separability hypothesis.","marker":"Minsky and Papert (1972)"},{"why":"Recent companion critique of the coherence of NLP research that the paper extends by defining core notions and capability assumptions.","marker":"Yogatama et al. (2019)"},{"why":"Provides the falsifiability criterion used to rank how precise and contestable a task's claimed capability involvement is.","marker":"Popper (1934)"},{"why":"Provides standard inter-annotator agreement methodology, adopted for the paper's dataset verification step.","marker":"Artstein and Poesio (2008)"},{"why":"Example of probing a trained model for syntactic structure, used to discuss how model-internal structure bears on capability claims.","marker":"Hewitt and Manning (2019)"},{"why":"List of desiderata for AGI environments, which the paper adapts into criteria for evaluating micro-worlds.","marker":"Adams et al. (2012)"}],"fun_headline_variants":["NLP tasks need to state which language skills they target","New datasets must specify the language ability they test","Task progress hinges on linking tasks to language competence","Without a capability link, new NLP tasks are vacuous","Explicit capability ties turn datasets into progress"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on there being a well-defined set of capabilities of a competent language user that can be broken into subsets tied to individual tasks; if that decomposition does not exist, evaluating tasks by their claimed capability links has no foundation.","fun_headline_variants_meta":{"raw":{"variants":["NLP tasks need to state which language skills they target","New datasets must specify the language ability they test","Task progress hinges on linking tasks to language competence","Without a capability link, new NLP tasks are vacuous","Explicit capability ties turn datasets into progress"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1306,"prompt_tokens":856,"completion_tokens":450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":376}},"tokens_in":472,"tokens_out":450,"duration_ms":4354,"temperature":1.0,"reasoning_tokens":376,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:34:36.935931+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a large battery of diverse language tasks with human participants and factor-analyze the performance matrix. If no stable task-specific capability components emerge — that is, a single general language factor explains nearly all variance — then the paper's assumption that the overall capability set CL decomposes into task-linked subsets CT loses its empirical grounding.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced the SHAPES dataset to push compositionality, an example of task improvement aimed at a specific capability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Historical notion of micro-worlds and the project of knitting them together; the paper uses this to frame environments and the separability hypothesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the falsifiability criterion used to rank how precise and contestable a task's claimed capability involvement is."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides standard inter-annotator agreement methodology, adopted for the paper's dataset verification step."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Example of probing a trained model for syntactic structure, used to discuss how model-internal structure bears on capability claims."},{"cited_title":"Storrs Hall, Alexei Samsonovich, Matthias Scheutz, Matthew Schlesinger, Stuart C","cited_arxiv_id":null,"evidence_quote":"List of desiderata for AGI environments, which the paper adapts into criteria for evaluating micro-worlds."}],"review_version":1}