{"id":"fa5b38a0-773e-421b-a3fa-895615dc65fc","arxiv_id":"2505.00603","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4 showed high recall but low precision in analogical matching for strategic problems, while humans showed the reverse, implying complementary roles for AI and human judgment.","lead":"This paper tested whether GPT-4 can match human analogical reasoning on business problems with two competing analogy sources. It found GPT-4 retrieves more analogies but picks wrong ones more often, while humans choose fewer but better analogies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4's low precision in the hint condition may be a response-format artifact: it lists both candidate stories, so per-claim scoring mechanically yields recall 1 and precision ~0.5. A forced-choice re-test is needed before claiming a human matching advantage.","rationale":"The reader's weakest assumption (non-independence of AI runs) is valid and should be fixed, but it is secondary: even with independent samples, the current confusion-matrix analysis would still not establish the central claim because the unit of analysis conflates response style with matching ability. The key observation is that GPT-4's hint-condition confusion matrices (Table 5) are almost exactly what one would obtain from a model that always names both stories: TP=15 and FP around 15 in each row. Counting each mention separately then forces recall=1 and precision approximately 0.5. Humans, who are less verbose and more selective in self-report, will mechanically show higher precision and lower recall. The Discussion's conclusion that the matching stage is the locus where humans retain a decisive advantage therefore rests on a measure that has not been validated against a single-best-match decision. A forced-choice experiment is a direct, inexpensive way to test whether the precision gap reflects genuine adjudication ability or merely a difference in how many candidates each agent volunteers. If the gap persists under forced choice, the paper's theoretical story about human causal-mapping superiority is supported. If it disappears, the paper's contribution reduces to a prompt-engineering and response-format finding. Given the strength of the claim ('decisive advantage', 'enduring comparative advantage'), the current evidence is insufficient; conditional acceptance with a request for this re-test is the right outcome.","tokens_in":16520,"tokens_out":11947,"duration_ms":126324,"concrete_test":"Run a forced-choice variant of the hint condition (Section 3.3) with both humans (n approximately 100) and GPT-4 (n at least 60, temperature above 0, seeds recorded): present the two source stories and one target problem, then ask 'Which of the two stories, if either, is the more appropriate analogy to this problem? Choose exactly one.' Score each response as a single selection. Compare selection accuracy, precision, and recall across agents. If GPT-4's accuracy is statistically indistinguishable from humans or higher, the original low-precision result is an artifact of open-ended generation and the matching-advantage claim fails; if humans still choose correctly more often, the claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that humans have a decisive advantage at the matching stage is not entailed by the current dependent measure. In Section 4.3, precision and recall are computed from confusion matrices in which each claimed story in a response is scored as a separate binary classification (Tables 5-6). Under the hint condition, GPT-4's counts are consistent with nearly every response naming both the correct and the incorrect story: AI: Radiation has TP=15, FP=14 and AI: Dolphin has TP=15, FP=15. Because each run is assigned to one target problem, these numbers imply that in both the Factory and HR conditions the model almost always mentioned both the relevant and the distractor story. When a system always lists both candidates, recall is definitionally 1 and precision approaches 0.5 regardless of any matching ability. Human participants, by contrast, typically report one story or none, which automatically yields higher precision and lower recall. The claimed human advantage in matching may therefore reflect differences in response format, verbosity, or instruction-following (the paper itself invokes a 'demand effect' in Section 4.3) rather than a superior capacity to adjudicate structurally appropriate analogues. The paper also reports no significance tests for the precision differences, and the AI runs may not be independent (Section 3.4 states AI output showed 'little variation'), compounding the problem. Without a forced-choice or single-best-match protocol, the headline conclusion is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares analogical reasoning in business contexts between human participants (n = 199) and GPT-4 runs (n = 60) using an experimental design with two source stories (radiation problem; dolphins/survivorship bias) and two target problems (factory logistics; HR pilot study). It measures correct analogy use via DAG-based coding and computes precision, recall, F1, and accuracy from confusion matrices in no-hint and hint conditions. The paper reports that, in the hint condition, GPT-4 shows high recall but low precision, while humans show high precision but low recall, and concludes that the matching stage is where humans retain a decisive advantage over LLMs. The paper also reports an error-type analysis distinguishing surface- from structure-based misapplications and discusses implications for human–AI complementarity in strategic decision making.","tokens_in":16860,"tokens_out":5541,"duration_ms":50898,"significance":"If the reported results were robust, the paper would make a useful contribution by extending the classical analogical-transfer paradigm to a two-source, two-target matching setting in naturalistic business vignettes, and by providing one of the first direct human–LLM comparisons on the full retrieval–mapping–evaluation pipeline. The pre-specification of correct analogies via DAGs, the inclusion of error-type coding, and the transparent discussion of limitations and the pre-registered attention check are strengths. However, the measurement and inference issues described below currently prevent the headline claim from being supported, so the paper's significance depends on whether those issues can be resolved in revision.","major_comments":[{"comment":"The central claim of a human matching advantage in the hint condition is not supported by the confusion-matrix counts because of a response-format confound. The AI counts (TP = 15, FP = 14 for Radiation; TP = 15, FP = 15 for Dolphin) are exactly what one would observe if every GPT-4 run listed both candidate source stories in its response; the per-claim scoring scheme used in the paper would then mechanically assign recall = 1.00 and precision ≈ 0.50 regardless of any real ability to distinguish structural from surface matches. Human participants, who typically report one story or none, automatically exhibit higher precision and lower recall under the same scoring. The paper acknowledges a 'demand effect' in §4.3 but does not control for it. A forced-choice protocol, or a scoring rule that considers only the participant's single best-supported match (or a top-1 precision metric), is needed before the data can be said to show that humans are better at matching.","section":"§4.3, Tables 5 and 6"},{"comment":"The AI sample cannot be treated as 15 independent observations per condition. The paper states in §3.4 that there was 'little variation in the AI output to the prompt' and that 15 observations per condition were therefore deemed sufficient. If the model's sampling is near-deterministic at the chosen settings, the 15 runs are not independent draws, and the reported precision/recall values (e.g., 0.52/1.00 for Radiation with hint) may reflect a single response pattern rather than a stable property of the model. The manuscript does not report the temperature or sampling parameters, nor the number of distinct response patterns per cell. The authors should report the effective number of unique responses and, if necessary, run additional trials with higher stochasticity or treat the model as a fixed effect in the analysis.","section":"§3.4, Tables 5 and 6"},{"comment":"The precision/recall comparisons are presented without any statistical tests or confidence intervals. Claims such as a 'decisive advantage' in the Discussion (§5) are not supported by formal inference. The paper should add tests appropriate to the data structure (each participant responds to one target problem under one hint condition, and AI runs are non-independent as noted above), or at minimum report exact binomial confidence intervals for the proportions underlying precision and recall. Without this, it is unclear whether the observed differences are within sampling variation.","section":"§4.3, Tables 4 and 6"},{"comment":"The abstract and Discussion characterize GPT-4 as having 'high recall' as a general property, but the perfect recall (1.00) is observed only in the hint condition; in the no-hint condition, GPT-4's recall is 0.00 for both stories (Table 4). Table 7 makes this condition-dependence explicit. The paper should qualify the high-recall claim to the hint condition, or explain why the no-hint behavior should be disregarded for the theoretical argument about matching. Otherwise the claim as stated is misleading.","section":"Abstract and §4.3, Table 7"}],"minor_comments":[{"comment":"There is a typo: 'we man that' should read 'we mean that'.","section":"§3.4"},{"comment":"Reporting 0/0 as zero is statistically questionable; it would be more appropriate to report the metric as undefined or to exclude that cell from the summary.","section":"Table 4 note"},{"comment":"The paper claims that intercoder reliability checks were used, but no reliability statistic (e.g., Cohen's kappa) is reported for the coding of correct analogical transfer or for the surface/structural error classification; please provide those values.","section":"§5, Limitations"},{"comment":"The Miller and Lin (2015) reference appears twice with identical details; the duplicate should be removed.","section":"References"},{"comment":"The y-axis label contains a typo: 'Distriubution' should be 'Distribution'.","section":"Figure 3"},{"comment":"The comparison of solvability with Gick and Holyoak (1983) is purely descriptive; consider adding a formal test or at least a caveat that the materials differ substantially between the studies.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and addresses a timely topic. The main concern is that the headline result—human advantage at matching—is currently vulnerable to a response-format artifact and to non-independent AI sampling. I believe the authors can address these issues by re-analyzing the existing data with a top-1 scoring rule (or by reporting the number of distinct response patterns) and by adding appropriate statistical inference. If the re-analysis does not support the claim, the conclusions should be substantially weakened. I recommend major revision rather than rejection because the design is promising and the data could plausibly support the claim after these fixes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The setup is genuinely new: extending Gick and Holyoak to a matching task with two sources and two targets, and running the same vignettes with humans and GPT-4. The error decomposition into surface versus structural misapplication is a useful lens, and the proposed division of labor (AI generates analogies, humans filter) is a natural implication if the asymmetry were real.\n\nBut the core empirical claim does not hold up under scrutiny. The stress-test note is right. In the hint condition, GPT-4 almost always mentions both source stories. Table 5 shows TP=15 and FP=14 or 15 for both stories. That means the model lists both candidates regardless of which is correct. When a system always lists both, recall is mechanically 1.0 and precision sits near 0.5. Human participants, by contrast, typically report one story or none. So the precision/recall gap is telling us about response format and the model's over-inclusive style, not about its capacity to adjudicate structural match. The paper itself invokes a \"demand effect,\" but that is an explanation, not a test.\n\nOther soft spots: no significance tests on the precision or recall differences; the AI sample is effectively very small because the 15 runs per cell show \"little variation in output\" and may be near-identical responses; and the abstract's high-recall claim applies only to the hint condition—without a hint, GPT-4 used no analogies at all. That no-hint result is interesting (LLMs apparently do not spontaneously retrieve analogies), but it is not about matching.\n\nWhat the paper does well: the careful construction of two structurally distinct source–target pairs, the DAG-based coding scheme, and the honest presentation of the confusion matrices. The limitations section is transparent about sample and task simplicity, though it misses the response-format confound.\n\nIs this worth a serious referee? Yes. The research question is important and the design has potential. But the current manuscript's headline conclusion is not supported. The fix is feasible: a forced-choice follow-up where the model must pick the single best analogy, or a scoring method that does not reward listing both. With that re-test, the asymmetry might still show up; right now, it is an artifact.\n\nMy recommendation: invite revision conditional on a forced-choice experiment, not accept the current evidence.","headline":"The two-source/two-target design is a welcome extension, but the headline human-matching advantage is almost certainly a response-format artifact: GPT-4 lists both stories, so recall is mechanically 1.0 and precision sits near 0.5.","tokens_in":17314,"tokens_out":3059,"would_cite":false,"duration_ms":31959,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Humans remain better than GPT-4 at the matching step of analogical reasoning, where choosing among plausible analogies requires causal structure, not just retrieval.","keywords":["analogical reasoning","strategic decision making","large language models","GPT-4","matching problem","precision and recall","causal mapping","human-AI complementarity"],"falsifier":"Re-run GPT-4 on the same two-story protocol for many more trials across varied sampling settings and random seeds, then inspect the distribution of claimed analogies in each condition. If the high-recall/low-precision pattern is stable across seeds, the paper's account is supported; if the outputs collapse to a few repeated texts or flip sharply with settings, the 15-run cells were not independent observations and the reported metrics overstate confidence in the model's behavior.","tokens_in":16356,"feed_emoji":"🤖","tokens_out":7918,"duration_ms":73179,"temperature":0.7,"pith_summary":"This paper asks whether GPT-4 can do the full job of analogical reasoning in strategic decisions—not just retrieve candidate analogies but pick the one whose causal structure fits the problem. In an experiment where two source stories (a converging-radiation technique and a survivorship-bias tale about dolphins) are matched against two business problems, GPT-4 retrieves nearly every applicable analogy but also applies many wrong ones, while human participants select fewer analogies and almost always select the right one. The paper concludes that the matching stage—adjudicating among superficially plausible analogues—is where human strategic reasoners keep a decisive advantage over current large language models. If correct, the practical consequence is a division of labor: language models as broad analogy generators, humans as evaluative filters.","feed_headline":"Humans still beat GPT-4 at picking the right analogy","feed_subtitle":"GPT-4 retrieves all analogies but applies many wrong ones; humans choose fewer, better matches.","key_machinery":"The load-bearing instrument is the experimental design itself. It turns the classic one-source-to-one-target analogy experiment into a matching problem: two source stories (the split-and-converge radiation technique and the dolphins-as-survivorship-bias story) pair with two target problems (city factory logistics and an HR training pilot), with known correct mappings from story to problem. Because the correct mappings are known in advance, every response can be placed in a confusion matrix and scored for precision, recall, F1, and accuracy. Directed acyclic graphs of each story's causal schema serve as the coding standard for whether an analogy was applied correctly, and the resulting errors are classified as surface-level or structural. This setup separates retrieval (did the agent use any analogy?) from matching (was it the structurally right one?), which is what allows the paper to locate the human advantage in the evaluative phase.","core_discovery":"On its own terms, the paper's central discovery is that human and machine analogical reasoning fail in mirror-image ways, and that the decisive bottleneck is the evaluative matching step. With an explicit hint that one of the two stories is relevant, GPT-4 achieved perfect recall (1.00) for both stories but precision of only 0.52 and 0.50, whereas humans achieved precision of 0.67 and 0.75 with recall of only 0.26 and 0.39. Without a hint, GPT-4 never spontaneously used the correct analogy, while humans did, with perfect precision but low recall. Coding of the wrong analogies shows that GPT-4's errors are driven mainly by surface similarity, such as linking the dolphin story to a coastal factory through the sea, whereas human errors come from applying an incorrect causal schema. The paper argues that this makes matching a first-order component of strategic analogical reasoning, and grounds a human-in-the-loop division of labor.","pith_inferences":["If GPT-4's output is as near-deterministic as the paper suggests, the perfect recall under the hint may partly reflect the model complying with the cue rather than genuine retrieval breadth; testing with withheld hints and more distractors would separate these.","Scaling the design from two sources to dozens could erode the human precision advantage, since evaluative capacity is itself limited; the proposed division of labor may then need a second AI filter before human review.","The paper's error classification relies on written justifications; asking both agents to justify why they rejected the non-chosen source would test whether the surface-versus-structural asymmetry is in reasoning or only in post-hoc explanation.","A direct implication for future model design is that the most promising lever is causal-schema discrimination, not wider retrieval."],"forward_implications":["In time-sensitive settings where missing a valid analogy is costlier than chasing false leads, GPT-4's perfect recall makes it a useful first-pass generator of candidate analogies.","In high-stakes, hard-to-reverse strategic choices, the human role as evaluative filter should be preserved, because human precision is markedly higher.","Because adding a choice among sources lowers solvability relative to one-to-one analogy tasks, performance claims from single-source analogy studies overstate what either humans or LLMs can do in realistic multi-source settings.","As LLMs expand the candidate set, training managers in structural comparison and causal mapping becomes more valuable rather than less.","The complementary error profiles (surface-driven for LLMs, causal-schema for humans) suggest ensembles that combine both agents could reduce overall error, though the paper does not identify the optimal ensemble design."],"supporting_citations":[{"why":"Supplies the radiation problem and the no-hint/hint protocol that the matching experiment extends.","marker":"Gick & Holyoak (1980)"},{"why":"Provides the one-to-one schema induction benchmark whose solvability rates the paper uses to show the cost of matching complexity.","marker":"Gick & Holyoak (1983)"},{"why":"Gives the structure-mapping account and the surface-versus-structural similarity distinction used to classify errors.","marker":"Gentner (1983)"},{"why":"Defines surface and structural similarity and the criterion for what counts as correct analogical transfer.","marker":"Holyoak & Koh (1987)"},{"why":"Frames mapping as constraint satisfaction, supporting the claim that matching is a distinct evaluative step.","marker":"Holyoak & Thagard (1989)"},{"why":"The MAC/FAC model separates retrieval from mapping/evaluation and underlies the paper's two-stage framing.","marker":"Forbus, Gentner, & Law (1995)"},{"why":"Review that distinguishes retrieval from mapping and motivates treating the matching/evaluation stage as underdeveloped.","marker":"Gentner & Smith (2013)"},{"why":"The demand-effect account the paper uses to explain GPT-4's over-application of analogies once a hint is given.","marker":"Gui & Toubia (2023)"}],"fun_headline_variants":["GPT-4 recalls too much, humans recall too little in analogies","Matching is the missing skill in AI analogical reasoning","Humans win on precision, GPT-4 wins on recall","GPT-4 picks surface matches, humans pick deep ones","AI over-retrieves analogies, humans over-filter them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on treating each of the 60 GPT-4 runs as an independent observation, but the paper reports little variation in the model's output and deems 15 runs per condition sufficient; if the sampling is near-deterministic, the reported precision and recall could describe one response pattern rather than stable properties of the model.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 recalls too much, humans recall too little in analogies","Matching is the missing skill in AI analogical reasoning","Humans win on precision, GPT-4 wins on recall","GPT-4 picks surface matches, humans pick deep ones","AI over-retrieves analogies, humans over-filter them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001269,"raw_usage":{"total_tokens":5187,"prompt_tokens":934,"completion_tokens":4253,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":4167}},"tokens_in":550,"tokens_out":4253,"duration_ms":32954,"temperature":1.0,"reasoning_tokens":4167,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:38:27.190012+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run GPT-4 on the same two-story protocol for many more trials across varied sampling settings and random seeds, then inspect the distribution of claimed analogies in each condition. If the high-recall/low-precision pattern is stable across seeds, the paper's account is supported; if the outputs collapse to a few repeated texts or flip sharply with settings, the 15-run cells were not independent observations and the reported metrics overstate confidence in the model's behavior.","supporting_citations":[],"review_version":1}