{"id":"02193763-c0f6-4934-bb65-cfa552b5e16c","arxiv_id":"2603.03414","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"AI's patchy performance stems from under-measured hidden cognitive processes, called cognitive dark matter, which could be made visible by collecting process-tracing and neural-behavioral data at scale.","lead":"This paper argues that AI systems fall short in subtle human abilities like self-checking and flexibility because we lack training data that reveals the hidden mental processes behind human behavior. It proposes a research program to collect exactly those data — from cognitive models, eye-tracking and think-aloud protocols, and brain recordings — to train more complete, less 'jagged' AI.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The transfer premise—that process-tracing and neural data can be converted into training signals that instill CDM abilities—is untested, and the cited evidence only covers perceptual/semantic domains, not the L3 functions central to the claim.","rationale":"The reader identified the same load-bearing concern: the assumption that cognitive-process measurements will transfer to AI training for L3 abilities. The paper is a research program, not a demonstrated result, and it explicitly hedges that the data may only benefit cognitive science. However, the central claim—that jaggedness stems from missing measurement and can be fixed by collecting these data—requires the transfer premise. Without a pilot demonstration, the claim remains speculative. The reader's CONDITIONAL verdict is appropriate: the framing and measurement-gap analyses are valuable, but acceptance of the causal thesis should be conditional on empirical transfer evidence. I see no reason to escalate to REJECT, as the paper does not overclaim proven transfer; it proposes a concrete plan. The proposed test would directly resolve whether the weakest assumption holds.","tokens_in":14649,"tokens_out":3902,"duration_ms":44638,"concrete_test":"Run a controlled pilot on a CDM-loaded task family, e.g., novel rule-induction puzzles like ARC-AGI-2. Collect think-aloud and eye-tracking data from humans solving a set of training puzzles. Fine-tune a base LLM on: (A) behavior-only (input/output pairs) and (B) behavior plus process traces (think-aloud transcripts, fixation sequences, self-correction events) as auxiliary supervision. Evaluate both on a disjoint held-out set of novel puzzles matched in difficulty and requiring the same cognitive flexibility/metacognition. Pre-register a minimum improvement threshold for (B) over (A) on held-out accuracy; if not met, the transfer premise is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim requires that the three proposed data types (latent variables from cognitive models, process-tracing like eye-tracking/think-aloud, paired neural-behavioral data) constitute causally efficacious training signals that improve generalization on CDM-loaded functions. The paper provides no derivation, pilot, or mechanism for this transfer. Process supervision (ref 18) rewards correct intermediate steps, which is not the same as cognitive process; brain-tuning studies (refs 28–34) improve perceptual/semantic robustness, a different regime from instilling metacognition, flexibility, or theory-of-mind. The authors explicitly admit they do not have scaling laws to predict how much neural data would be needed for direct training. Thus, even if the measurement gap is real (as the surveys suggest), the causal jump from 'under-measured' to 'can be remedied by collecting these data' is unsupported. This is the load-bearing assumption; if it fails, the jaggedness explanation reduces to a correlation with no actionable mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes that the jagged capability landscape of modern AI systems is caused by a missing training signal, termed 'cognitive dark matter' (CDM): cognitive functions that shape human behavior but are hard to infer from observable behavior alone (metacognition, cognitive flexibility, lifelong learning, abductive reasoning, social/emotional reasoning). It argues that current AI benchmarks and large-scale neuroscience datasets both under-represent these L3 functions, and it proposes a research program based on three data types—latent variables from cognitive models, process-tracing data, and paired neural-behavioral data—to train AI on cognitive process rather than outcome. Three small empirical analyses are presented: a chess-app example, a survey of benchmark tiers in model release documents, and a survey of neuroimaging/literature coverage. The central claim is that better measurement of CDM will produce more general, less jagged AI, with the same data also advancing cognitive neuroscience.","tokens_in":1509,"tokens_out":1422,"duration_ms":44557,"significance":"If the transfer premise held, the paper would provide an actionable and fairly original research program that connects AI benchmarking, cognitive modeling, and large-scale neuroscience. The dual-benefit framing is attractive, and the authors are explicit that scaling laws for neural-data training are currently unknown. The paper also makes a falsifiable prediction: collecting the proposed data types should improve performance on CDM-loaded tasks beyond what outcome-only data of the same scale would provide. However, the empirical support is exploratory: the three analyses are small, partially manual, and not accompanied by inter-rater reliability or confidence intervals, and the central causal claim rests on an untested transfer assumption. The paper's strength is as a synthesis and proposal, not as a demonstration.","major_comments":[{"comment":"The load-bearing claim is that process-tracing data, neural recordings, and cognitive-model latents can serve as training signals that instill metacognition, flexibility, and social reasoning into AI. No pilot, derivation, or mechanism is provided. The cited brain-tuning studies (refs 28–34) concern perceptual/semantic robustness, not L3 functions; process supervision (ref 18) rewards correct intermediate steps, which is not the same as latent cognitive process. The paper itself admits the absence of scaling laws. This transfer premise needs at least a proof-of-concept or a concrete falsifiable test, e.g., comparing CDM-enriched versus outcome-only supervision on the same L3 benchmarks with matched compute.","section":"CDM, Jagged Intelligence, and human–AI interactions; 'What to collect next'"},{"comment":"The chess example is presented as evidence of missing metacognition, but it is based on N=3 models × 10 runs, with manual scoring and no inter-rater reliability, confidence intervals, or quantified failure rates. The interactive Claude Code sessions are described qualitatively as 'failed to elicit it.' As an illustrative anecdote this is acceptable, but as evidence for the core claim it is under-powered. Report the exact counts, scoring rubric, and inter-rater agreement, or explicitly label it as anecdotal.","section":"Methods: Chess endgame evaluation"},{"comment":"The categorization of benchmarks and neuroimaging studies into the Liu et al. L1/L2/L3 taxonomy is subjective and mostly manual, with LLM-assisted tagging for some datasets. The 50 GPT-5.2 runs for benchmark classification are reported as 'stability' but no agreement statistic is given. Without inter-rater reliability or a validation of the LLM labels, the quantitative 'gap' shown in Figures 2 and 3 is suggestive but not established. The qualitative conclusion may survive, but the numerical presentation overstates precision.","section":"Methods: AI benchmark analysis and neuroimaging surveys; Figures 2–3"},{"comment":"The paper asserts that under-measurement *causes* jagged performance, but it does not rule out alternative explanations: L3 tasks may be harder because they require more data in general, different architectures, or better world models, independent of whether the missing signal is 'cognitive process.' The constant-hazard-rate analysis of Ord (ref 9) suggests subtask success rates p and effective step count n; improving p could occur through many mechanisms, not specifically CDM-enriched data. The central hypothesis should be stated with a concrete discriminating experiment, otherwise it risks being unfalsifiable.","section":"CDM, Jagged Intelligence, and human–AI interactions"}],"minor_comments":[{"comment":"Typographical errors: 'isjagged', 'musthave', and missing spaces before citations in a few places. Figure 3 labels 'Claude 4.5' while the text and Methods say 'Claude Opus 4.5'.","section":"Abstract and body text"},{"comment":"The figure caption does not define what counts as 'intensive neuroimaging datasets' or report the number of papers per source, making the y-axis scale hard to interpret. Please clarify the counts and sources.","section":"Figure 2 caption"},{"comment":"The Methods say 'GPT-5.2 with high reasoning effort' was used to assign cognitive functions, but no details are given on the prompt, temperature, or how the 50 runs were aggregated. If the analysis is to be reproducible, these details are needed.","section":"Methods: AI benchmark analysis"},{"comment":"The '~500 hours' milestone for large-scale datasets is stated without a citation or quantitative justification. It is a reasonable heuristic but should be framed as such.","section":"What to collect next"},{"comment":"Reference [43] is an incomplete author list with 'K. Allen ... et al.'; [54] is a test instrument rather than a peer-reviewed article. Please verify the reference metadata.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-suited to a perspective/opinion venue, but the current framing as an empirical claim without supporting evidence may be problematic. The reader's stress-test concern about the transfer premise is valid and lands on the central causal claim. A major revision that explicitly reframes the paper as a proposal with falsifiable predictions, or that adds a small proof-of-concept study, would strengthen it substantially. The empirical analyses, as presented, should either be reported with proper uncertainty/agreement or labeled as exploratory."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What to know: this is a research program proposal, not a proven result. The core claim is that jagged AI performance stems from a missing training signal—cognitive processes that shape behavior but aren't visible in behavior alone—and that collecting process-level data (latent variables from cognitive models, eye-tracking/think-aloud, paired neural-behavioral data) would fill the gap. The paper is honest about that: it says outright that scaling laws are unknown and that the program might miss the mark. That honesty earns credit.\n\nThe best part is the synthesis. Jaggedness, process supervision, brain-tuning, and the L1/L2/L3 taxonomy previously existed as separate threads. Putting them under one explanatory frame—and pointing out that both AI benchmarks and large-scale neuroimaging datasets skew away from L3 abilities—is genuinely useful. The two surveys, though rough, suggest the same picture: benchmarks cluster on L2, intensive imaging on L1, and L3 is scarce in both. That observation alone is worth citing. The chess example is anecdotal but effective as an illustration of a metacognitive failure.\n\nSoft spots are real and proportionate. The load-bearing causal assumption—that process-tracing and neural data will transfer to L3 abilities like metacognition or flexibility—is untested. The cited brain-tuning work is mostly about perceptual and semantic robustness, which is not the same regime. The analyses are exploratory: the chess probe is N=3 models, manually scored without confidence intervals; the benchmark classification uses one GPT-5.2 run repeated fifty times but reports no inter-rater agreement; the neuroimaging survey uses LLM tagging without validation. The paper acknowledges these weaknesses, but the reader should treat the empirical sections as pilot observations, not evidence.\n\nStill, the central logic is not circular. The paper proposes a hypothesis and a data-collection agenda; it doesn't fit parameters to hide a tautology. The L1/L2/L3 taxonomy is borrowed, but the interpretive claim about what it means for AI training is new. I think the causation claim needs a pilot demonstration—even a small one—before the program is fully convincing, but the framing itself is sound enough to take seriously.\n\nWho benefits: people thinking about AI training data, benchmarking strategy, and the neuroscience-for-AI pipeline. It makes a good reading-group piece because it will provoke useful disagreement. I'd cite it if I wrote about benchmark design or brain-inspired training. Send it to peer review as a perspective/agenda paper, with referees who will press the authors on the transfer mechanism. It doesn't deserve a desk reject.","headline":"A wide-framing hypothesis paper that names a real pattern and a data agenda; the empirical support is thin, but the idea is worth taking seriously.","tokens_in":15388,"tokens_out":1391,"would_cite":true,"duration_ms":17691,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the jagged intelligence of AI systems stems from a missing training signal—cognitive dark matter—and that collecting process-level data from human minds and brains can supply that signal.","keywords":["cognitive dark matter","jagged intelligence","metacognition","cognitive flexibility","neural-behavioral data","process tracing","AI benchmarks","cognitive process training"],"falsifier":"A controlled comparison: train matched models on the same CDM-heavy tasks with and without process-tracing or neural supervision, then test on novel L3 tasks such as rule-switch flexibility and confidence calibration. If process-supervised models show no improvement over behavior-only baselines at comparable data scale, the thesis that the missing signal is latent process data is falsified.","tokens_in":14467,"feed_emoji":"🧠","tokens_out":4945,"duration_ms":48297,"temperature":0.7,"pith_summary":"The paper proposes that the uneven, 'jagged' intelligence of modern AI—adept at some tasks, surprisingly poor at others—is not a fixed ceiling but a measurement gap. The missing ingredient is cognitive dark matter: brain functions such as metacognition, cognitive flexibility, social reasoning, and emotional intelligence that shape behavior yet are hard to infer from behavior alone. The paper surveys current AI benchmarks and large-scale neuroscience datasets and finds both are skewed toward capabilities AI has already mastered, while these hard-to-measure functions are largely absent. To close the gap, it proposes collecting three kinds of data—latent variables from cognitive models, process-tracing data such as eye-tracking and think-aloud protocols, and paired neural-behavioral data—so models can be trained on cognitive process rather than behavioral outcome alone. If this is right, the path to more general, less jagged AI runs through measuring the mind, with better understanding of human cognition as a dual payoff.","feed_headline":"AI gaps trace to a missing signal: cognitive dark matter","feed_subtitle":"Paper says collecting process-level human data could train models on cognition, not just outcomes.","key_machinery":"The central object is cognitive dark matter (CDM), defined as brain functions that meaningfully shape behavior yet are hard to infer from behavior alone. The argument runs through a three-tier taxonomy used in the paper (L1: mastered capabilities like vision and language; L2: partial progress; L3: rarely explored CDM-loaded functions such as cognitive flexibility, social reasoning, and metacognition). The proposed mechanism is a trio of data types: latent variables sampled from large-scale cognitive models, process-tracing data (eye-tracking, mouse-tracking, think-aloud), and paired neural-behavioral data. These are meant to supply the missing training signal by exposing the causal cognitive","core_discovery":"On the paper's own terms, the central claim is that the jagged intelligence landscape of AI systems arises from a missing training signal, not from lack of scale or raw capability. CDM-loaded functions are largely unmeasured in current AI benchmarks and in large-scale neuroscience datasets, so models never receive the signal needed to acquire them. The paper argues that making hidden cognitive processes visible—through latent variables, process-tracing data, and paired neural-behavioral data—would let AI train on process rather than outcome, producing models that generalize more smoothly and fail in human-legible ways. The authors present the measurement gap as both an opportunity and a warn","pith_inferences":["An implicit prediction is that returns from process-level data should show up specifically on L3 capabilities; a scaling-law-style curve for eye-tracking or neural data against flexibility and calibration metrics would test this.","The chess-endgame case suggests a concrete diagnostic: measuring whether a model abandons a failing strategy after feedback could serve as a quantitative metacognition score before and after CDM-style training.","The thesis connects jaggedness to safety: chronic miscalibration, hallucination, and perseveration may share a single missing metacognitive signal, so better measurement could improve reliability as well as capability.","If the latent signals turn out to be epiphenomenal, the paper's own dual-benefit framing still justifies the dataset program, but the AI-transfer claim would need a different mechanism."],"forward_implications":["Benchmarking practice would shift: frontier model evaluation suites should include L3-heavy tests, because current suites systematically under-test the functions where jaggedness hides.","The bottleneck for less jagged AI becomes data collection, not architecture: large, intensive datasets for metacognition, flexibility, social reasoning, and related functions become a prerequisite.","Training on process data should make failures more human-legible—detectable and interpretable—because models would learn the cognitive processes that generate human outputs.","Process supervision alone is insufficient for most tasks because correct intermediate steps are unknown; the paper's three data types are meant to supply those latent processes at scale.","A dual benefit follows: even if AI transfer falls short, the datasets would be foundational for understanding human cognition."],"fun_headline_variants":["AI's blind spot: cognitive dark matter","To fix AI, make hidden cognition measurable","Cognitive dark matter explains AI's jagged gaps"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that hidden cognitive processes—metacognitive states, attention, neural activity—when added to training data are causally efficacious signals that transfer human-like capabilities to models; if they are only correlates of behavior, training on them may not close the jaggedness gap.","fun_headline_variants_meta":{"raw":{"variants":["AI's blind spot: cognitive dark matter","To fix AI, make hidden cognition measurable","Cognitive dark matter explains AI's jagged gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000728,"raw_usage":{"total_tokens":3072,"prompt_tokens":696,"completion_tokens":2376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":2331}},"tokens_in":440,"tokens_out":2376,"duration_ms":15600,"temperature":1.0,"reasoning_tokens":2331,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T02:32:43.313806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison: train matched models on the same CDM-heavy tasks with and without process-tracing or neural supervision, then test on novel L3 tasks such as rule-switch flexibility and confidence calibration. If process-supervised models show no improvement over behavior-only baselines at comparable data scale, the thesis that the missing signal is latent process data is falsified.","supporting_citations":[],"review_version":1}