{"id":"346ee71f-d38e-4e56-82d5-a7c632da56d2","arxiv_id":"2411.15626","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A perspective paper maps how humans and machines generalize differently and argues that aligning these generalization behaviors is essential for human-AI teaming.","lead":"AI alignment usually focuses on goals and values, but this perspective argues that how machines generalize, versus how humans generalize, is a key missing piece. The paper maps definitions, methods, and evaluation of generalization across cognitive science and AI, and lists open challenges for human-AI teaming.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central argument's load-bearing premise is normative—that human-like generalisation is the right alignment target—and the paper's own output-level alignment caveat (§2) undercuts the inference from process/product differences to a need for new methods.","rationale":"The reader's weakest assumption correctly identifies the normative premise as the most load-bearing point. My stress-test adds the observation that the paper's own output-level alignment caveat in §2 makes the inference in §3.4 weaker than it appears: process/product differences do not automatically create an alignment requirement. A substitution test would show whether the argument depends on 'human-like' at all. I do not recommend changing the UNVERDICTED verdict: this is a perspective paper, and the concern is a flagged caveat about scope rather than a detectable factual error. The taxonomy and survey content are valuable, and the central research agenda should be read as conditional on the normative premise being accepted.","tokens_in":24435,"tokens_out":4912,"duration_ms":46796,"concrete_test":"Analytical substitution test: rewrite the conclusion of §3.4 twice, replacing 'human-like generalisation ability' with (a) 'robust task competence under distribution shift' and (b) 'the user's own stated generalisation preferences.' If the argument 'different processes/products ⇒ different generalisation ⇒ need new methods' survives both substitutions unchanged, then the paper's specific appeal to human-like generalisation is not doing argumentative work and the central claim reduces to a tautology; if it does not survive, identify the precise property of human generalisation that makes the inference valid and check whether the paper supplies evidence for that property (it currently does not).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §3.4 is conditional: if we wish to align machines to human-like generalisation ability, new methods are needed. The conditional is nearly tautological; the load-bearing content is the antecedent, i.e., the normative choice of human-like generalisation as the alignment target. The paper asserts this choice but never defends it against alternatives such as task-level competence under distribution shift, user-specified generalisation preferences, or calibrated robustness. This matters because the paper itself restricts alignment to outputs: §2 says 'it is sufficient that the output of a learnt model is aligned with human cognitive concepts' and footnote 1 disclaims that capabilities enabling preferences need to be similar between humans and AI. With that restriction, the fact that humans and machines use different processes and produce different products (abstractions vs. distributions) does not by itself imply a need for new methods; what would need to be shown is that current alignment methods leave machine generalisation behaviour misaligned with human judgements on novel inputs in a way that feeds back into preference satisfaction. The examples given (hallucinations, adversarial sensitivity, OOD brittleness) are suggestive but are not tied to a formal criterion of 'generalisation alignment.' In addition, 'human-like generalisation' is not a single coherent benchmark: the paper acknowledges human overgeneralisation (stereotyping) and context-dependence, so the target itself needs specification. This is a normative/argumentative gap rather than an internal inconsistency, and for a perspective paper it is the main reason the result is unverdictable rather than wrong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This perspective paper argues that human and machine generalisation differ systematically and that aligning machine generalisation with human generalisation is an overlooked aspect of AI alignment. It develops a trichotomy of generalisation as process, product, and operator; maps AI method families (statistical, knowledge-informed, and instance-based) onto these notions; and surveys evaluation practices and open challenges. The central claim, stated in Section 3.4, is that because humans and machines use different processes and produce different products, and if one wishes to align machines to human-like generalisation ability, then new methods are needed. The paper does not present new experiments or formal results; its contribution is a synthesis and a research agenda for interdisciplinary work at the intersection of AI and cognitive science.","tokens_in":24755,"tokens_out":7733,"duration_ms":66672,"significance":"If the central claim is accepted, the paper identifies a genuine gap: alignment research has focused on preferences, while generalisation behaviour has received less attention. The paper's strengths are its interdisciplinary scope, the clear process/product/operator taxonomy, and the structured comparison of statistical, analytical, and instance-based methods in Tables 2 and 3. It is also honest about caveats such as human overgeneralisation and context-dependent categorisation. However, as a position paper it does not provide a formal definition of generalisation alignment, a defended normative target, or quantitative evidence for the capability attributions in Table 3. The paper would be substantially strengthened if it specified a measurable criterion for generalisation alignment and made explicit which claims are consensus views and which are the authors' own position. Its value is heuristic: it maps a territory and poses questions rather than settling them.","major_comments":[{"comment":"The central claim is conditional: 'if we wish to align machines to human-like generalisation ability, we need new methods to achieve machine generalisation.' As stated, this is close to tautological, and the load-bearing content is the normative antecedent. The paper does not defend human-like generalisation as the appropriate alignment target against alternatives such as task-level competence under distribution shift, user-specified generalisation preferences, or calibrated robustness. Because the abstract and Section 1 frame alignment as acting according to preferences, the relationship between preference alignment and generalisation alignment must be made explicit. Please add a substantive defence of the normative premise, or explicitly reframe the contribution as a conditional research agenda.","section":"§3.4"},{"comment":"The output-level sufficiency caveat undercuts the inference from process/product differences to the need for new methods. Section 2 states that for effective human-AI teaming it is sufficient that the output of a learnt model is aligned with human cognitive concepts, and footnote 1 disclaims that the capabilities enabling preferences need to be similar between humans and AI. If only outputs matter, the fact that humans and machines use different processes and produce different products does not by itself imply that current alignment methods are inadequate. The paper should show a concrete way in which machine generalisation behaviour on novel inputs feeds back into preference satisfaction, and it should reconcile Section 2 with Section 6, where realignment is said to impose stricter requirements for collaboration at the process level.","section":"§2"},{"comment":"No operational definition of 'generalisation alignment' is provided. Section 5.4 lists three aspects (distributional shifts, under- and overgeneralisation, and memorisation versus generalisation) but does not specify a criterion or metric that would distinguish aligned from misaligned generalisation. Consequently, the examples offered in Sections 1 and 4.1 (hallucinations, adversarial sensitivity, OOD brittleness) are suggestive but do not substantiate a claim of misalignment. Moreover, 'human-like generalisation' is not a single coherent benchmark: the paper acknowledges human overgeneralisation (stereotyping) and context-dependent categorisation, so the target needs to specify which human generalisation is intended (expert versus lay, normative versus descriptive, individual versus population). A measurable definition is needed to make the central claim testable.","section":"§5.4"},{"comment":"Table 3's binary '+'/'-' assignments are too coarse and are in places inconsistent with the body text. For example, Section 4.1 states that few-shot and zero-shot mechanisms allow statistical methods to build on learned representations, yet Table 3 assigns 'learning from a few samples' a '-' for statistical methods; Section 4.2 discusses recent progress in compositional generalisation in deep learning through analytical components, yet 'compositionality' is '-' for statistical methods. The table also lacks quantitative evidence or specific references in the evaluation column. Please either soften the table to reflect the caveats discussed in Section 4, or provide evidence for each entry.","section":"Table 3"}],"minor_comments":[{"comment":"There is a typographical error in the sentence 'machines are increasingly tested for their ability tohandle complex data'; it should be 'to handle complex data'.","section":"§5"},{"comment":"The phrase 'diluting generalisability to align with empirical observations' is unclear and should be rephrased.","section":"§6"},{"comment":"Reference [75] is the arXiv preprint of the present manuscript; in a journal version this self-citation should be replaced with the published version or removed.","section":"References"},{"comment":"The table header 'Method EvaluationS A I' is malformed; the layout should be cleaned up so that the method families and evaluation column are clearly separated.","section":"Table 3"},{"comment":"The 'operator' notion could be defined more sharply: as presented, applying a generalisation to new data is close to the standard notion of model prediction, and the distinction between 'product' and 'operator' should be clarified.","section":"§3.3"},{"comment":"The definition of overgeneralisation as making 'over-confidently false predictions' is unusual; in cognitive science overgeneralisation typically means applying a rule or concept too broadly. Please align the terminology with the cognitive-science literature or justify the deviation.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"This is a broad community position paper arising from a Dagstuhl seminar. The main scientific risk is that the central claim is conditional and near-tautological, so the authors should be pushed to commit to a measurable definition of generalisation alignment and to defend the normative choice of human-like generalisation as the target. The self-citation of the arXiv preprint as reference [75] should be removed. The breadth of authors and references is a strength, but the paper should distinguish established results from the authors' own research agenda."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a solid perspective paper, not a research result. The best part is the organising trichotomy — generalisation as process, product, operator — and the Table 3 comparison of statistical, analytical, and instance-based methods. It gives a shared vocabulary that alignment discussion could use. The survey of historical and current work is broad, well-cited, and the evaluation section on under- and overgeneralisation, memorisation, and distribution shifts is sensible.\n\nThe soft spot is the central claim in §3.4. The authors say that because humans and machines use different processes and produce different products, they generalise differently, and if we wish to align machines to human-like generalisation ability we need new methods. The conditional is fine but nearly tautological — the real weight is on the antecedent, 'if we wish to align machines to human-like generalisation ability.' That target is never defended. The paper's own §2 says output alignment is sufficient and process similarity is not required (footnote 1). If that is the standard, process/product differences alone do not imply a need for new methods. You would need to show that current alignment methods leave machine outputs misaligned with human judgements specifically on novel, out-of-distribution inputs, and that this affects preference satisfaction. The examples — hallucinations, adversarial attacks, OOD brittleness — are suggestive but not tied to a criterion for generalisation alignment. So the headline conclusion is a research agenda, not a derivation.\n\nTable 3 is another caveat. Strict '+' / '−' categorical entries without quantitative backing are handy for a bird's-eye view but overclaim. Statistical methods are marked '−' for learning from a few samples, despite few-shot and in-context learning existing (though admittedly not human-like); instance-based methods get '+' for OOD robustness, but only given a suitable representation, which the paper itself notes in §4.3. That is a limitation, not fatal, if the table is read as an idealised summary.\n\nSelf-citation is heavy, but this is a synthesis by people central to neurosymbolic and human-centric AI; it is not a circular derivation. It deserves a real review. I would send it to peer review with a request to either defend the normative target or soften the claim to an explicitly conditional agenda. It is not a must-read, but it is a useful reference for anyone writing about generalisation and alignment.","headline":"A useful conceptual map of human and machine generalisation, but the central call for new alignment methods rests on an unargued normative premise that the paper's own output-alignment caveat undermines.","tokens_in":25344,"tokens_out":3410,"would_cite":true,"duration_ms":30923,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This perspective paper argues that AI alignment must cover not only what machines should prefer but how they generalise, and it maps the gap between human and machine generalisation across three dimensions to show why current alignment is…","keywords":["generalisation","AI alignment","human-AI teaming","cognitive science","out-of-distribution generalisation","neurosymbolic AI","compositionality","evaluation"],"falsifier":"If human–AI teams performed at or above the best individual human or best AI across a broad battery of out-of-distribution and compositional tasks without any generalisation-specific alignment, the claim that alignment must target generalisation would be weakened. A direct test is to run paired human and model assessments on the same distribution-shift and compositionality benchmarks and check whether equalising preference agreement alone leaves systematic generalisation disagreement.","tokens_in":24287,"feed_emoji":"🧠","tokens_out":3576,"duration_ms":32361,"temperature":0.7,"pith_summary":"This paper argues that AI alignment is incomplete if it only makes AI act according to human preferences: it must also align how machines generalise. Humans and machines arrive at different products of generalisation — humans form sparse, compositional abstractions and causal concepts from few examples, whereas statistical AI fits probability distributions that stay close to the training data. The paper maps these differences across three dimensions — notions, methods, and evaluation of generalisation — and concludes that aligning machine generalisation to human-like ability requires new methods, not just new evaluation. A sympathetic reader would take this as a research agenda: treat generalisation behaviour as a first-class alignment target in human-AI teaming.","feed_headline":"Align AI's generalisation, not just its preferences","feed_subtitle":"Perspective paper argues human-AI teaming requires matching machine generalisation to human abstraction, composition, and causal learning.","key_machinery":"The organising device is a three-part decomposition of generalisation: as a process (abstraction, extension, analogy), as a product (categories, rules, prototypes, exemplars, probability distributions), and as an operator (applying a product to new data). This decomposition shows that process-level differences drive product-level differences, and that the operator-level divergence — how far a model can be applied beyond its data — is the alignment gap. The paper also categorises machine methods by their source-target relationship: statistical generalisation transfers observations to a population, knowledge-informed generalisation seeks evidence for an explicit theory, and instance-based generalisation translates from specific cases to new specific cases.","core_discovery":"The central claim is that because humans and machines use different processes (e.g., abstraction vs. data-driven learning), they arrive at different products (e.g., categories and rules vs. probability distributions) that generalise differently; therefore, if we wish to align machines to human-like generalisation ability as an operator, we need new methods to achieve machine generalisation. The paper substantiates this by surveying three families of AI methods — statistical, knowledge-informed, and instance-based — showing they have complementary strengths: statistical methods offer universal approximation but black-box behaviour and poor out-of-distribution performance, knowledge-informed methods provide compositionality and explainability but are limited to formalisable domains, and instance-based methods are robust to distribution shift and support lifelong learning but depend critically on the chosen representation. On evaluation, it argues that current practice — train/test splits, distribution-shift measures, and contamination checks — is necessary but insufficient because human notions of under- and overgeneralisation involve adapting to task variations and avoiding hallucination-like overconfidence beyond the data.","pith_inferences":["If human-like generalisation becomes the alignment target, then benchmarks for alignment should be built around human performance on distribution shifts and compositional tasks, not just human preference ratings, which would make alignment measurable across the three dimensions the paper defines.","A concrete test suggested by the framework: measure whether a model's error profile on out-of-distribution tasks matches the error profile of human raters on the same tasks; a systematic mismatch would quantify misalignment in operator-level generalisation.","The paper leaves implicit that aligning generalisation may require giving AI systems explicit causal models, common-sense priors, or compositional inductive biases rather than more data, because the human advantage draws on prior-driven abstraction rather than raw statistical learning.","The argument extends naturally to generative AI: hallucination is an overgeneralisation rather than a preference failure, so alignment methods that only tune preferences may miss the generalisation errors that most undermine trust in human-AI teams."],"forward_implications":["Alignment research should target generalisation behaviour, not only stated preferences, because human-AI teaming fails when machine generalisation behaviour diverges from human expectation.","No single AI method family currently provides human-like generalisation across the board: statistical methods lack out-of-distribution generalisation and compositionality, knowledge-informed methods are limited to formalisable domains, and instance-based methods depend on suitable representations.","Evaluation of generalisation must go beyond IID train/test splits to measure distributional shifts, under- and overgeneralisation, and the memorisation–generalisation boundary, since test-set leakage in foundation models can invalidate reported performance.","Neurosymbolic and hybrid methods, plus new theory for few-shot and zero-shot feasibility, are needed to close the operator-level gap between humans and machines.","The framework implies that generalisation guarantees, such as compositionality and robustness bounds, should be part of the certification of AI systems in human-AI teaming contexts."],"supporting_citations":[{"why":"Defines alignment as 'make AI systems act according to our preferences', the target the paper extends.","marker":"[26]"},{"why":"No-free-lunch theorem grounds the claim that statistical generalisation is bounded without prior knowledge.","marker":"[168]"},{"why":"Fodor and Pylyshyn's critique of connectionism supplies the compositionality challenge that separates statistical ML from human cognition.","marker":"[49]"},{"why":"Tenenbaum et al. provide the cognitive-science account of human generalisation as Bayesian structure learning and abstraction.","marker":"[151]"},{"why":"PAC framework is the formal theory of generalisation as an operator under IID assumptions.","marker":"[142]"},{"why":"Survey of hallucination in natural language generation exemplifies overgeneralisation as an alignment failure.","marker":"[78]"},{"why":"Li and Flanigan's finding of task contamination in ChatGPT shows how evaluation can be invalidated, undermining presumed generalisation.","marker":"[100]"},{"why":"Meta-analysis showing human-AI teams lag behind the best individual provides the motivation for the alignment problem.","marker":"[155]"},{"why":"Neurosymbolic AI is cited as the promising bridge that combines statistical, knowledge-informed, and instance-based strengths.","marker":"[69]"}],"fun_headline_variants":["Align AI's generalisation, not just its preferences","Generalisation gap: the overlooked key to AI alignment","To align AI, align its generalisation habits","Humans and machines generalise differently; align that","AI alignment needs matching human and machine generalisation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that human-like generalisation is the right standard for AI alignment; the paper asserts this as the goal but does not defend it against alternatives such as task-level competence or value alignment without cognitive mimicry.","fun_headline_variants_meta":{"raw":{"variants":["Align AI's generalisation, not just its preferences","Generalisation gap: the overlooked key to AI alignment","To align AI, align its generalisation habits","Humans and machines generalise differently; align that","AI alignment needs matching human and machine generalisation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1345,"prompt_tokens":947,"completion_tokens":398,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":563,"tokens_out":398,"duration_ms":4428,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:04:36.805205+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If human–AI teams performed at or above the best individual human or best AI across a broad battery of out-of-distribution and compositional tasks without any generalisation-specific alignment, the claim that alignment must target generalisation would be weakened. A direct test is to run paired human and model assessments on the same distribution-shift and compositionality benchmarks and check whether equalising preference agreement alone leaves systematic generalisation disagreement.","supporting_citations":[{"cited_title":"How to grow a mind: Statistics, structure, and abstraction","cited_arxiv_id":null,"evidence_quote":"Tenenbaum et al. provide the cognitive-science account of human generalisation as Bayesian structure learning and abstraction."},{"cited_title":"When combi- nations of humans and ai are useful: A systematic review and meta-analysis","cited_arxiv_id":null,"evidence_quote":"Meta-analysis showing human-AI teams lag behind the best individual provides the motivation for the alignment problem."}],"review_version":1}