{"id":"0f5288a3-3ec6-47b4-9544-99d904e71d87","arxiv_id":"2607.08274","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a real UK law-enforcement deployment, crime analysts selectively used AI linkage predictions and validated them against behavioural matrices, attending most to MO similarity and geography.","lead":"Analysts used an AI crime-linkage tool selectively, cross-checking predictions against traditional behavioural evidence rather than trusting scores alone. The in-situ mixed-methods study shows how explanations and non-AI matrices must be tightly integrated for high-stakes policing tools to be usable and trusted.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged sample and design limits; the multi-modal evidence still supports the observational claim.","rationale":"The paper's strongest claim is carefully scoped as an industrial mixed-methods observation of how SCAS analysts actually engaged with LATIS predictions and explanations. Multi-modal data (observation notes, USE/UXIV/QUIS scores, fixation percentages/heatmaps, matrix open counts and durations) converge on selective use plus traditional validation. The threats the reader lists are real and correctly motivate CONDITIONAL rather than full ACCEPT, yet they are already disclosed and do not falsify the reported behaviours within the studied setting. No deeper load-bearing flaw (e.g., circular measurement, mis-specified AOIs, or contradiction between modalities) is present once the faulty radar-plot component is discarded. Hence the reader's verdict and confidence stand; no adjustment is required.","tokens_in":16996,"tokens_out":497,"duration_ms":6213,"concrete_test":"Re-analyse the existing eye- and mouse-tracking logs after excluding the first (most complex) session for every participant; if the rank-order of fixations (VA_ID > MO > geo) and the high matrix-open frequency remain unchanged, the learning-order and complexity confounds do not drive the headline interaction patterns.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is observational (selective use of AI scores + frequent cross-check against the behavioural matrix + attention to MO/geo features). The reader's weakest assumption already captures the main threats: n=6 volunteers, fixed order, three pre-selected series with a true link forced into the top-20, and discarded radar-plot data (Sections 3.1, 3.3, 6). These constrain generalisability and exact replication but do not invert the reported interaction patterns. Eye-tracking (Figs. 5–6) shows consistent fixation on VA_ID, MO similarity and geographical proximity across sessions; mouse-tracking (Table 1) shows the matrix opened ≥10 times in 12/16 sessions, almost always for single-case comparison; surveys and notes corroborate selective validation rather than oracle use. No internal contradiction appears once the radar-plot results are set aside. The claim therefore holds under the paper's own scope (in-situ usability evidence for this tool and team).","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports a mixed-methods industrial usability evaluation of LATIS, an AI decision-support tool for behavioural crime linkage co-developed with the UK NCA SCAS. Six analysts completed 16 sessions on three real crime series (top-20 ranked lists seeded with at least one true link), combining direct observation, eye-tracking, mouse-tracking of the behavioural matrix, and post-session USE/UXIV/QUIS surveys plus free-text. Findings indicate selective use of AI probability scores, frequent cross-validation against the non-AI behavioural matrix, attention to all five model features (highest on MO similarity and geographical proximity), positive ease-of-use/ease-of-learning ratings but mixed usefulness, and concrete suggestions for better integration of explanations with traditional analytic practices.","tokens_in":17258,"tokens_out":1212,"duration_ms":24850,"significance":"If the observational patterns hold, this is a useful contribution to HCI and AI-for-policing: it supplies rare in-situ multi-modal evidence (eye- and mouse-tracking plus surveys) from expert analysts working with real sensitive data, rather than mock tasks or self-report alone. The triangulation supports design implications that AI predictions should be presented with feature-level explanations aligned to domain priorities (MO, geography) and tightly coupled to non-AI evidence (behavioural matrix) to support verification and partial trust. Strengths include the co-design history, operational setting, transparent handling of the radar-plot data error, and an explicit threats-to-validity section. These elements make the work more actionable than typical lab studies of XAI in high-stakes domains.","major_comments":[{"comment":"Section 3.1 (and Threats 6): The sample comprises only six volunteer SCAS analysts (three analysts, three senior) who completed a fixed-order sequence of three pre-selected series, each engineered so that at least one true link appears in the top-20. While the multi-modal data consistently show selective AI use and matrix cross-checking within this sample, the design risks volunteer bias, learning effects (explicitly noted in rising USE scores across sessions), and inflated engagement from the forced-link seeding. These factors are load-bearing for the general claims about “how analysts use AI” and the derived design implications; the paper should either (a) more tightly scope all claims to “this team and tool under these conditions” or (b) add a short sensitivity discussion quantifying how the forced-link and order constraints might affect the observed validation rates.","section":"Section 3.1 / Section 6"},{"comment":"Section 3.3 and Findings 4.2–4.3: The radar-plot component was discarded after participants correctly identified incorrect data values. This is handled transparently, yet the tool description (Section 2.2, Figure 1) and study-design overview still present the radar plot as a core explanation view. Consequently the RQ2 claim that “analysts attended to … the model features presented as explanations” rests solely on the ranked-list colour cells and the behavioural matrix. The manuscript should explicitly restate that the attention and valuation findings apply only to the ranked-list feature explanations and matrix, and remove or clearly flag any residual implication that the radar plot contributed usable evidence.","section":"Section 3.3 / Section 4.2"},{"comment":"Section 4.2 (Figures 5–6) and RQ2: Fixation percentages and heat-maps are purely descriptive; no inferential statistics, confidence intervals, or baseline comparisons are reported for differences among the five features or across sessions. Given that the paper’s second research question asks whether analysts attend to all features and that the strongest attention claim is used to argue alignment with model importance, at least non-parametric tests or bootstrapped intervals on the fixation proportions would make the “greatest attention to MO and geographical proximity” statement more robust. Without them the claim remains impressionistic.","section":"Section 4.2"}],"minor_comments":[{"comment":"Section 5: typographical error “final descisions remain the responsibility”.","section":"Section 5"},{"comment":"Figure 5 caption: missing space in “Behavioural Matrixrepresents the button”.","section":"Figure 5"},{"comment":"Table 1: the “Time (min)” columns are hard to parse; consider separating average open duration from total open time more clearly, and note that two participants completed only two sessions.","section":"Table 1"},{"comment":"Section 4.1: the statement that overall USE scores “increased from session 1 to session 2, and then increased further in session 3” would benefit from reporting the actual mean scores or a simple plot so readers can judge the magnitude of the learning effect.","section":"Section 4.1"},{"comment":"References: a few DOIs and arXiv links appear incomplete or point to future versions; double-check consistency before camera-ready.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The work is a solid, carefully executed industrial HCI study under genuine operational constraints (secure site, real crime data, eye-tracker logistics). The small-n and design limitations are real but typical for this setting; once the three major points are addressed by tighter scoping and modest additional analysis, the paper should be publishable. No concerns about novelty disclosure or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first in-situ evaluation of a real AI crime-linkage tool (LATIS) with SCAS analysts, using real sensitive series, eye-tracking, mouse-tracking, observation and standard questionnaires. That combination is new; prior work is either algorithmic with no user study or interview/survey only.\n\nWhat it does well is straightforward. Across 16 sessions the multi-modal data line up: analysts fixated most on VA_ID, MO similarity and geographical proximity (Figs 5–6), opened the behavioural matrix often (Table 1: ≥10 times in 12/16 sessions, almost always single-case), and treated AI scores as decision support rather than oracles. Surveys and notes match. The design implications—keep feature explanations, integrate the traditional matrix tightly, allow filtering—are concrete and useful for HCI/XAI people building high-stakes tools. Threats section is honest about learning effects, fixed order and the discarded radar plot.\n\nSoft spots are real but already flagged and proportionate. n=6 volunteers, three pre-selected series with a true link forced into the top-20, fixed task order, and one broken component. That limits generalisability and exact replication; it does not invert the observed interaction patterns. No circularity, no invented entities, citations look appropriate (tool background papers plus the relevant policing-AI and usability literature). Math is not the point; the observational claims are grounded in the recorded fixations, clicks and scores.\n\nThis is for people designing or evaluating AI decision support in policing or other high-stakes expert domains. It will not reshape a field, but it is the kind of careful industrial evidence that is still scarce. A serious editor should send it to referees; the sample-size and design limits are revision material, not desk-reject material. I would bring it to reading group and would cite the interaction findings when writing about XAI adoption in law enforcement.","headline":"Solid first industrial multi-modal usability study of an operational AI crime-linkage tool; the selective-use and feature-attention findings hold under the paper’s own scope despite the small volunteer sample.","tokens_in":17840,"tokens_out":481,"would_cite":true,"duration_ms":5702,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Analysts use AI crime-linkage suggestions selectively and still verify them against traditional behavioural evidence.","keywords":["Artificial Intelligence","crime linkage","decision making","usability study","explainable AI","human-AI interaction","eye tracking","law enforcement"],"falsifier":"A larger, non-volunteer sample of analysts working unselected live cases with no planted top-twenty link who rarely open the behavioural matrix and accept AI ranks without cross-check would overturn the selective-validation claim.","tokens_in":17936,"feed_emoji":"🔍","tokens_out":873,"duration_ms":19179,"temperature":0.7,"pith_summary":"This industrial study asks how specialist crime analysts actually work with an AI tool that ranks offences most likely linked to an index case and shows the features driving those ranks. In a real law-enforcement setting, with real series and real analysts, people engaged with the tool, attended to the feature explanations, and valued having them, yet routinely opened a non-AI behavioural matrix to cross-check candidates rather than treating the scores as final. Ease of use and ease of learning scored high; perceived usefulness was more mixed, and analysts asked for better ways to mark reviewed cases and hide features or columns that did not matter for the decision at hand. The paper’s central point is that high-stakes AI decision support must tightly integrate model explanations with the methods analysts already trust, and that only in-situ evaluation with operational users and real data can show how that integration works.","feed_headline":"Analysts verify AI crime links against old methods","feed_subtitle":"In-situ study finds selective trust, heavy attention to MO and geography, and demand for tighter workflow integration.","key_machinery":"A mixed-methods usability evaluation of the LATIS crime-linkage interface: direct observation, eye-tracking of fixations on ranked scores and feature cells, mouse-tracking of behavioural-matrix openings, and post-session usability questionnaires, run with operational analysts on three real crime series.","core_discovery":"Analysts used the AI predictions selectively and frequently validated them against behavioural (non-AI) evidence, reflecting partial trust and continued reliance on established practice. They attended to all presented model features, with the heaviest attention on MO similarity and geographical proximity, valued those explanations, and still opened the behavioural matrix repeatedly to verify ranked candidates. The tool is therefore used as decision support that must be checked, not as a substitute for traditional analysis.","pith_inferences":["The same selective-trust pattern is likely in other high-stakes domains where experts already have strong non-AI methods they must defend.","Guaranteeing a true link in the top twenty may have inflated attention to AI ranks relative to noisier real deployments.","Continued fixation on MO and geography lower in the list suggests analysts use those feature columns as a secondary filter once probability bars fade.","Without usable multi-feature comparison views, reliance on the behavioural matrix for verification may stay higher than necessary."],"forward_implications":["AI crime-linkage tools should surface feature-level explanations beside ranked scores because analysts attend to and value them.","Traditional non-AI checks such as the behavioural matrix must be embedded tightly in the same workflow so verification is efficient.","In-situ evaluation with real users and real data is required to surface selective trust and integration needs before wider deployment.","Interaction flexibility (mark reviewed rows, hide irrelevant features or columns) is needed to keep focus and reduce effort.","Perceived usability can rise with familiarity alone, so early mixed usefulness scores do not by themselves rule out later operational acceptance."],"fun_headline_variants":["Analysts selectively verify AI crime links with behavioural evidence","Crime analysts check AI predictions against traditional MO methods","Partial trust: analysts validate AI linkages with non-AI data","Analysts lean on MO and geography while verifying AI crime ranks","AI crime tool used as checkable support not analysis substitute"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That six volunteer analysts completing sixteen sessions on three pre-selected series, each engineered so a true link appears in the top twenty, represent how crime-linkage units will use the tool in ordinary operational work.","fun_headline_variants_meta":{"raw":{"variants":["Analysts selectively verify AI crime links with behavioural evidence","Crime analysts check AI predictions against traditional MO methods","Partial trust: analysts validate AI linkages with non-AI data","Analysts lean on MO and geography while verifying AI crime ranks","AI crime tool used as checkable support not analysis substitute"]},"model":"grok-4.5","effort":"low","cost_usd":0.007568,"raw_usage":{"total_tokens":1826,"prompt_tokens":799,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":75680000,"prompt_tokens_details":{"text_tokens":799,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":945,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":799,"tokens_out":82,"duration_ms":8259,"temperature":1.0,"reasoning_tokens":945,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T10:13:11.678304+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger, non-volunteer sample of analysts working unselected live cases with no planted top-twenty link who rarely open the behavioural matrix and accept AI ranks without cross-check would overturn the selective-validation claim.","supporting_citations":[],"review_version":1}