{"id":"ae29875e-b9ec-4af2-a716-84fc2e9432df","arxiv_id":"2604.04333","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"West Point cadets display less automation bias and more calibrated trust in algorithmic advice than the general public in a target identification task.","lead":"The study runs a survey experiment where West Point cadets and a similar public group identify targets, receive advice from an algorithm or human, and can revise their answers. Results indicate cadets show better calibrated trust in AI advice than the general public.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"External validity of the simplified target-ID task to real military decisions remains untested and load-bearing.","rationale":"The reader's weakest assumption correctly isolates the mapping from lab measure to real-world military judgment. Full-text methods confirm the task is a standard vignette with no additional validity evidence, so the concern stands and keeps the verdict at UNVERDICTED.","tokens_in":1723,"tokens_out":301,"duration_ms":24480,"concrete_test":"Re-run the identical target-identification protocol with the same cadet cohort but add a high-stakes incentive (e.g., performance-tied bonus or simulated mission outcome) and compare effect sizes on advice adherence to the original low-stakes condition; if the cadet-public gap shrinks or reverses, the original claim does not generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (cadets show better-calibrated trust than the public sample) is measured via change in identification after receiving algorithmic or human advice in a low-stakes survey task. For this to support claims about reduced cognitive distortion in military AI use, the task must capture the same mechanisms that operate under operational time pressure, accountability, and lethal consequences. The paper reports no manipulation checks, no high-fidelity simulation arm, and no within-cadet analysis linking task behavior to actual training or deployment experience. Demographic matching does not address whether the observed difference is an artifact of the artificial setting rather than military education per se.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports results from a survey experiment with 236 US Military Academy cadets and a demographically matched public sample. Participants performed a target-identification task, received advice labeled as either algorithmic or human, and were allowed to revise their initial judgment. The central finding is that cadets exhibited better-calibrated trust (lower automation bias and algorithm aversion) than the public sample, suggesting that military education and AI exposure can reduce cognitive distortions in interactions with decision-support systems.","tokens_in":1822,"tokens_out":522,"duration_ms":24389,"significance":"If the result holds, the work provides rare empirical evidence on how military training shapes human-AI interaction in a domain with high stakes for international security. It directly addresses a gap in the literature on automation bias and algorithm aversion within professional military populations and offers a falsifiable claim that education can produce more calibrated reliance on algorithmic advice.","major_comments":[{"comment":"The headline claim that cadets show reduced cognitive distortion rests on the assumption that the low-stakes target-identification task elicits the same mechanisms that operate under operational time pressure, accountability, and lethal consequences. No manipulation checks, high-fidelity simulation arm, or within-cadet correlation with actual training/deployment experience are reported to support this mapping.","section":"Methods and Results (target identification task description)"},{"comment":"The abstract states that 'the findings are limited,' yet the manuscript provides no explicit discussion of how the artificial setting, absence of real-world consequences, or lack of demographic matching on military-specific variables (e.g., prior AI exposure, command experience) might artifactually produce the observed cadet-public difference.","section":"Abstract and Discussion"}],"minor_comments":[{"comment":"The sample size (N=236 cadets) and exact statistical tests, effect sizes, and confidence intervals for the key cadet-public comparison are not summarized in the abstract or early sections, making it difficult to assess the precision of the 'better calibrated trust' claim.","section":"Abstract"},{"comment":"The paper does not report whether the algorithmic advice was actually more accurate than human advice in the task, which is necessary to distinguish calibrated trust from simple accuracy following.","section":"Experimental design"}],"recommendation":"major_revision","confidential_remarks":"The external-validity concern raised in the stress-test note is load-bearing; without additional validation data or explicit scope limitations, the manuscript risks overclaiming policy relevance for real military AI deployment."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important considerations regarding the generalizability of our findings. We address each major comment below and have revised the manuscript accordingly where feasible.","responses":[{"response":"We acknowledge that the target identification task is a controlled, low-stakes survey experiment and does not replicate operational conditions such as time pressure or lethal consequences. The design was chosen to isolate the effects of advice source (algorithm vs. human) on judgment revision in a standardized manner across both samples, enabling a direct comparison of automation bias and algorithm aversion. No manipulation checks for perceived stakes or high-fidelity simulation arms were included, as the study was a survey-based experiment focused on population differences rather than ecological validity. We will add an expanded limitations subsection in the Discussion to explicitly address the assumptions required to map these results to high-stakes military contexts and note the absence of within-cadet correlations with training experience.","revision_made":"partial","referee_comment":"[Methods and Results (target identification task description)] The headline claim that cadets show reduced cognitive distortion rests on the assumption that the low-stakes target-identification task elicits the same mechanisms that operate under operational time pressure, accountability, and lethal consequences. No manipulation checks, high-fidelity simulation arm, or within-cadet correlation with actual training/deployment experience are reported to support this mapping."},{"response":"We agree that while the abstract notes the findings are limited, the Discussion would benefit from more explicit treatment of these potential artifacts. We will revise the Discussion to include a dedicated paragraph addressing how the artificial setting and lack of real-world consequences could influence results, as well as the incomplete matching on military-specific variables such as prior AI exposure and command experience. This will clarify possible alternative explanations for the observed differences without overstating generalizability.","revision_made":"yes","referee_comment":"[Abstract and Discussion] The abstract states that 'the findings are limited,' yet the manuscript provides no explicit discussion of how the artificial setting, absence of real-world consequences, or lack of demographic matching on military-specific variables (e.g., prior AI exposure, command experience) might artifactually produce the observed cadet-public difference."}],"tokens_in":1366,"tokens_out":509,"duration_ms":19886,"standing_objections":["We cannot add manipulation checks, a high-fidelity simulation arm, or within-cadet correlations with deployment experience, as these would require new data collection beyond the existing survey experiment."]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main result is that West Point cadets change their answers less extremely after receiving algorithmic advice than a demographically matched public sample does. They measure this through a pre-post design on a target identification task, which lets them track automation bias and algorithm aversion directly by comparing shifts after algo versus human advice. This is a clean application of an established survey method to a new and relevant group. The direct cadet-public comparison is the clearest addition here, and the abstract is straightforward about the directional finding and its boundaries. The setup itself is simple enough that the measurement logic is easy to follow. The soft spot is external validity, exactly as the stress-test flags. The task is low-stakes, no time pressure, no accountability, and no real consequences, so the observed difference may not travel to operational settings where decisions carry lethal weight. The paper reports no manipulation checks or high-fidelity arms to test whether the survey behavior tracks actual training or deployment experience. Demographic matching helps on the surface but does not address whether military education itself drives the gap or whether the artificial context does. The abstract already notes the findings are limited, which is accurate given the design. This is useful for people working on human-AI interaction in security studies or behavioral aspects of military decision-making. It is not yet strong enough on its own for broad claims about escalation risks or AI integration, but the comparison is worth referee time if the authors can add detail on methods, sample characteristics, and any robustness checks. I would send it to review with a request to strengthen the validity discussion rather than desk reject.","headline":"Cadets show better calibration than the public in a simple target-ID task, but the artificial setup and admitted limits keep the military-AI implications modest.","tokens_in":2295,"tokens_out":390,"would_cite":false,"duration_ms":17253,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Paper studies human-AI decision biases in military cadets vs public; RS framework derives physics from logic of distinction","alignment":"orthogonal","rationale":"The paper's central machinery is a within-subject survey experiment measuring switching rates after algorithmic/human advice in a target-ID task, testing automation bias and algorithm aversion. This has zero overlap with RS theorems (e.g., reality_from_one_distinction, AlexanderDuality.alexander_duality_circle_linking, Cost.Jcost_pos_of_ne_one, or any J-cost/phi-ladder/8-tick derivations). RS has no opinion on cognitive biases or military decision-making.","tokens_in":56658,"confidence":"high","tokens_out":148,"duration_ms":7009,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"West Point cadets display better calibrated trust in algorithmic advice than the general public in a target identification task.","keywords":["automation bias","algorithm aversion","military AI","decision support systems","West Point cadets","human-AI interaction","target identification","cognitive bias"],"falsifier":"A study observing actual military operators using AI decision support in field exercises or simulations to check if their trust calibration matches the cadet results.","tokens_in":2611,"feed_emoji":"🪖","tokens_out":571,"duration_ms":42438,"temperature":0.7,"pith_summary":"The paper examines how military training affects human interaction with AI decision support systems by comparing West Point cadets to a similar public sample. Participants completed a target identification task, received advice from either an algorithm or a human, and could update their judgments. Cadets showed more appropriate reliance on the algorithmic advice, avoiding both over-trust and undue skepticism that the public exhibited. This suggests that military education helps mitigate cognitive biases that could lead to errors in AI-assisted conflict decisions.","feed_headline":"West Point cadets calibrate AI trust better than public","feed_subtitle":"A survey experiment finds West Point cadets calibrate trust in AI advice more accurately than civilians.","key_machinery":"Survey experiment with a target identification task where participants receive advice from an algorithm or human analyst and have the chance to reassess their initial judgment.","core_discovery":"West Point cadets are less prone to cognitive distortion than members of the general public, displaying better calibrated trust in algorithmic decision support systems. The experiment directly measured changes in identification after receiving advice, revealing that cadets adjusted their assessments in line with the quality of the input more effectively than civilians.","pith_inferences":["Similar training approaches could be adapted for other high-stakes AI users like doctors or pilots to improve calibration.","These findings point to education as a lever for shaping AI's impact on international security.","Further studies could test if the effect holds in more complex, realistic military scenarios beyond the lab task."],"forward_implications":["Military personnel may be less likely to err due to automation bias or algorithm aversion when using AI in operational settings.","AI integration in militaries could be managed with lower risk of miscalculation if training emphasizes calibrated trust.","Exposure to AI through education influences how humans interact with decision support in high-stakes environments.","The role of human judgment in war may evolve differently in professional military forces than in civilian contexts."],"fun_headline_variants":["West Point cadets calibrate AI trust more accurately than civilians","Military cadets better calibrate trust in AI advice than public","West Point cadets adjust to AI advice more effectively than civilians","Cadets from USMA show better calibrated AI trust than civilians"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The target identification task and survey responses capture real-world susceptibility to automation bias and algorithm aversion in military decision-making.","fun_headline_variants_meta":{"raw":{"variants":["West Point cadets calibrate AI trust more accurately than civilians","Military cadets better calibrate trust in AI advice than public","West Point cadets adjust to AI advice more effectively than civilians","Cadets from USMA show better calibrated AI trust than civilians"]},"model":"grok-4.3","cost_usd":0.006859,"raw_usage":{"total_tokens":3100,"prompt_tokens":659,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":68590500,"prompt_tokens_details":{"text_tokens":659,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2377,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":659,"tokens_out":64,"duration_ms":27906,"temperature":1.0,"reasoning_tokens":2377,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T20:23:33.719955+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study observing actual military operators using AI decision support in field exercises or simulations to check if their trust calibration matches the cadet results.","supporting_citations":[],"review_version":1}