{"id":"73d2a100-2f22-4522-ac0b-be1dd8bf71fa","arxiv_id":"2506.00218","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using a novel survey elicitation method, the authors estimate that Finns consider false positive rates around 10 to 15 percent in travel pre-screening acceptable, with air travel tolerated more than sea travel.","lead":"A survey of 1550 Finnish adults finds that people tolerate surprisingly high numbers of false positives in algorithmic travel surveillance, roughly 10 to 15 percent of flagged passengers, and that they accept air-travel surveillance more than sea-travel surveillance. The study offers the first quantitative estimate of public tolerance for false positives in EU passenger name record systems, relevant to legal and policy debates about the base rate fallacy in mass surveillance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 9.7% and 14.7% legitimate false-positive rates are not identified by the data; they are artifacts of the arbitrary 200-passenger cap and an untested log-normal interval-regression assumption in §4.3.3.","rationale":"Read in good faith, the paper's core contribution is a new way to quantify a contextual-integrity acceptability threshold for opaque algorithmic surveillance. The qualitative result—Finns tolerate very high false-positive counts and tolerate air travel more than sea travel—is credible and internally consistent with the attitudinal items (Q3–Q12). The paper is also unusually candid about limitations, including normalcy bias, the absence of false negatives, and simplifying assumptions about human verification. However, the headline numbers in the abstract are load-bearing for the paper's practical implications: the discussion uses the 12.7% figure to compute roughly 140,000 manual checks per day at EU level. Those numbers inherit every weakness of the interval construction in §4.3.3. The binary response at a single F does not elicit a threshold; it only indicates which side of F the respondent's threshold lies. Treating that as evidence about all values up to 200 is a strong, untested assumption, and the log-normal distribution is similarly assumed without sensitivity analysis. The reader's weakest_assumption identifies exactly this issue. The proposed concrete test would settle it by re-estimating without the log-normal assumption and without the 200 cap: if the estimates are stable, the concern does not land and the paper's quantitative claims stand; if they move substantially, the abstract should present ranges or only the qualitative finding. In either case, the appropriate verdict remains conditional, so no adjustment to the reader's verdict is needed.","tokens_in":21371,"tokens_out":5842,"duration_ms":61181,"concrete_test":"Re-estimate the model in §4.3.3 under two perturbations: (a) replace the log-normal assumption with a log-logistic or nonparametric interval-censored estimator; (b) cap security intervals at 30 (the maximum F shown) instead of 200, treating larger values as unidentifiable. If the resulting median acceptable counts move outside the 9.7–14.7% range, or if the air/sea ordering changes, the headline rates are artifacts of the assumptions. A stronger direct test is a follow-up survey with an open-ended or adaptive maximum-acceptable-false-alarms question that does not anchor on 200.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3.3 converts each binary vignette response into a wide interval: security at F false positives becomes (F,200], privacy becomes [0,F). This transformation is the only source of the reported 'legitimate false positive count.' Two features do the real work. First, the 200 ceiling is the vignette's arbitrary denominator: a respondent who prioritizes security at F=30 is treated as accepting every count up to 200, and a respondent whose true acceptable count exceeds 200 is censored at 200. Second, a log-normal distribution is assumed for the latent count; the reported values (e.g., 25.83 false positives for air travel, Table 2) are exp(linear predictor), i.e., the median of that assumed distribution, divided by 200 to obtain a 'rate.' No diagnostic or robustness check is reported for either assumption. Because vignette F values only ran from 1 to 30 (§4.1.3), all estimates above 30 are extrapolation. The qualitative conclusion that Finns tolerate high false-positive counts is supported by the response pattern, but the abstract's precise 9.7%/14.7% rates (actually gender-specific extremes from Table 2, not overall estimates) are not directly measured and could shift under a reasonable re-specification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a nationally representative Finnish survey (N=1550) with randomized air-travel and sea-travel conditions, using a vignette that presents a binary choice between prioritizing security and prioritizing privacy at a given number of false positives. Responses are converted into intervals and analyzed with log-normal interval regression, from which the authors estimate 'legitimate' false-positive counts and rates. The paper claims that Finns accept very high false-positive rates (9.7% for sea and 14.7% for air in the abstract), that air travel is more acceptable than sea travel, and that these findings challenge purely rights-based justifications of algorithmic surveillance. The discussion connects the results to contextual integrity, the base-rate fallacy, and the limits of the EU legal framework.","tokens_in":21642,"tokens_out":3989,"duration_ms":42700,"significance":"If the quantitative estimates were credible, the paper would make a distinctive empirical contribution to the surveillance-privacy literature: it offers one of the first population-level measurements of acceptable false-positive rates for PNR-style pre-screening, and it explicitly frames the result as a challenge to individual-rights-based legitimacy arguments. The study has clear strengths: a large, stratified sample; randomized experimental conditions; a transparent questionnaire appendix in both English and Finnish; consistent between-condition differences in the attitudinal items; and an unusually honest limitations section that acknowledges normalcy bias and the human-in-the-loop simplification. The paper is also careful to distinguish the descriptive use of contextual integrity from a full normative evaluation. However, the headline false-positive rates are not directly elicited quantities; they are outputs of an interval regression whose two central assumptions—the 200-passenger ceiling and the log-normal latent distribution—are neither justified nor stress-tested.","major_comments":[{"comment":"The interval transformation in Section 4.3.3 is the sole source of the reported point estimates. A security-prioritizing response at F false positives is converted to the interval (F, 200], and a privacy response to [0, F), which makes the vignette's 200-passenger denominator an upper bound. The log-normal distribution is assumed without diagnostics, and the vignette only showed F values from 1 to 30 (§4.1.3), so the model's upper confidence limits (e.g., 33.03 in Table 2) and any estimate treated as a population threshold depend on extrapolation beyond the displayed range. The Table 2 rates are the exponentiated linear predictor—the median of the assumed log-normal distribution—divided by 200, not directly measured acceptability rates. Please report robustness checks (for example, alternative distributional assumptions, nonparametric interval bounds, a sensitivity analysis varying the vignette denominator, or at minimum the raw proportion choosing security at each F by condition) before these numbers are presented as legitimate false-positive rates.","section":"§4.3.3, Table 2"},{"comment":"The abstract's headline rates '9.7% and 14.7% for sea and air travel' do not match Table 2's combined estimates of 11.1% for sea and 12.9% for air; 9.7% and 14.7% are instead the gender-specific extremes (male/sea and female/air). The Discussion also refers to an air-travel rate of 12.7%, which differs from the Table 2 air-travel combined value of 12.9%. The paper should consistently report a single set of estimates, clearly distinguishing pooled estimates from gender-specific subgroups.","section":"Abstract, §5.3, §6"},{"comment":"The ANOVA tests used for model selection are nested comparisons of interval-regression models, so the conclusions that age and the number of true positives do not affect acceptance are conditional on the same interval/log-normal specification that produces the point estimates. Reporting raw associations—for example, the proportion of security-prioritizing responses by true-positive count and by age group—would provide a specification-free check of these null results and would strengthen the model-selection claim.","section":"§5.2"}],"minor_comments":[{"comment":"'Course-grained' should be 'coarse-grained'.","section":"§3.2"},{"comment":"'Rights-bases approaches' should be 'rights-based approaches'.","section":"§6"},{"comment":"Reference [69] spells the author's name as 'Chrstian Thönnes'; the correct spelling is 'Christian Thönnes'.","section":"References"},{"comment":"Reference [23] is dated 'Retrieved 21 December 2025', which postdates the arXiv submission date of the manuscript; please verify the retrieval date.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for FAccT and addresses a timely and important question. The core qualitative finding—that a representative sample of Finns accepts double-digit false-positive rates, with stronger acceptance for air than for sea travel—is plausibly supported by the raw response patterns and by the attitudinal items. The gating issue is the unvalidated interval-regression machinery behind the precise headline rates; this is fixable with additional robustness analyses and more careful reporting. I would not reject, but I would not accept until the estimation assumptions are defended or relaxed and the Abstract/Table 2/Discussion numbers are harmonized."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading for one reason: it gives the first nationally representative estimate of how much false positive error the public will tolerate in EU PNR-style travel surveillance. The Finnish data (N=1550) and the air-versus-sea comparison are new, and the qualitative finding — Finns tolerate double-digit false positive rates, and tolerate them more for air than sea travel — is reasonably well supported. The method is a novel application of the authors' collective criticism technique to contextual integrity, which is a genuine empirical contribution even if the underlying method is not new. Credit where due: the survey design is careful, the limitations section is honest, and the attitudinal questions corroborate the vignette results.\n\nThe soft spots are real, though not fatal. The precise headline rates (9.7% and 14.7% in the abstract) are not direct measurements. They come out of an interval regression that converts each binary vignette response into an interval capped at the vignette's 200 passengers, then fits a log-normal distribution to those intervals. The 200 cap and the log-normal assumption are both arbitrary and untested, and the paper reports no robustness checks. A respondent who prioritizes security at 30 false positives is treated as accepting everything up to 200 — that's a modeling choice, not a fact. And the abstract is worse than imprecise: it presents the lowest and highest gender-specific rates (male/sea and female/air) as if they were the overall sea and air estimates. Table 2's actual overall rates are 11.1% and 12.9%. That should be corrected.\n\nThat said, the stress-test's claim that the rates are \"artifacts\" goes too far. The raw response pattern — a majority choosing security even at high false positive counts — is directly observable, and the significance tests for travel mode and gender are standard and fine. The qualitative conclusion that Finns are highly tolerant of false positives, and more tolerant in air travel, does not depend on the log-normal assumption. What depends on the assumption is the exact number, and that number is being overstated in the abstract.\n\nThe paper would be stronger with a sensitivity analysis (different caps, maybe a nonparametric description of the response curve) and with data/code released. As it stands, it is a solid empirical contribution for people working on surveillance, contextual integrity, and EU data protection law. It deserves a serious referee, and I'd want to see it after revision. I would cite it for the new data.","headline":"New Finnish survey data on public tolerance for false positives in travel surveillance is genuinely useful, though the headline false positive rates are model-dependent and the abstract misreports gender-specific numbers as overall rates.","tokens_in":22137,"tokens_out":2930,"would_cite":true,"duration_ms":28165,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A nationally representative survey finds Finns accept double-digit false-positive rates in algorithmic travel pre-screening, with air travel more accepted than sea travel.","keywords":["privacy trade-off","surveillance","contextual integrity","PNR Directive","false positives","data protection","public security","algorithmic surveillance"],"falsifier":"Run the same vignette with a different denominator, for example 500 or 1,000 flagged passengers, and also ask respondents directly for the maximum acceptable number of false alarms; if the estimate does not scale with the denominator, or direct answers diverge from the regression-based threshold, the reported rates are an artifact of the 200-passenger frame and the assumed curve.","tokens_in":21197,"feed_emoji":"✈️","tokens_out":9997,"duration_ms":90868,"temperature":0.7,"pith_summary":"This paper asks how much false-positive error the public will tolerate in algorithmic passenger pre-screening, a statistical feature of the EU's passenger-name-record (PNR) surveillance system that is hidden from passengers by law. Using a nationally representative survey of 1,550 Finnish adults, it finds that respondents judged double-digit false-positive rates legitimate under a contextual-integrity standard, roughly 11-13% combined for sea and air travel, and between 9.7% and 14.7% depending on travel mode and gender. It also finds air travel surveillance consistently more acceptable than sea travel, and women accept more false positives than men. The authors argue this bottom-up measure of legitimacy complements top-down rights-based EU regulation, which has struggled with the base-rate fallacy. If the finding holds, critiques of mass surveillance that focus on individual rights violations may not match what the public actually accepts.","feed_headline":"Finns accept up to 15% false alarms in travel screening","feed_subtitle":"Representative survey finds air-travel pre-screening error rates of 9.7-14.7% are accepted; sea travel less so.","key_machinery":"The central device is a vignette-based trade-off task paired with interval regression. Respondents see 200 flagged passengers with a random number of true and false positives, choose between prioritizing security or protecting the privacy of innocent passengers, and that binary choice is converted into an interval on the acceptable number of false positives, for example choosing security at 10 false positives means the acceptable count lies in the interval (10, 200]. Assuming the acceptable count follows a log-normal distribution, the paper fits interval regression models to estimate the mean acceptable false-positive count, which it then reports as a false-positive rate. The collective-criticism origin of the method allows nearby vignette values to inform the same estimate.","core_discovery":"Working within the contextual-integrity framework, the paper claims that the Finnish public's perception of when algorithmic travel surveillance becomes illegitimate can be quantified as a false-positive threshold, and that this threshold is much higher than rights-based critiques would predict. In a vignette where 200 passengers are flagged and a randomly sampled number turn out to be false alarms, respondents chose between prioritising security and protecting innocent passengers' privacy; interval regression on those binary choices yields estimates of 22.10 acceptable false positives out of 200 (11.1%) for sea travel and 25.83 (12.9%) for air travel when pooling genders, with the gender-specific figures spanning 9.7% to 14.7%. The estimates are significantly higher for air than for sea travel, and women's thresholds are about 29% higher than men's. The authors take this as evidence that the contextual integrity of the information flow is not breached by high false-positive counts, and that acceptability tracks the travel context rather than the identical rights violations suffered by the flagged passengers.","pith_inferences":["Going beyond the paper, the same vignette run in a lower-trust or more ethnically diverse population would test whether the double-digit thresholds are a Finnish high-trust result; the paper itself flags this cultural-context caveat.","Because the vignette fixes the flagged pool at 200 passengers, the estimates may be anchored to that denominator; varying the pool size in a replication would show whether the 10-15% rates are a framing artifact.","If many respondents who chose security would have accepted more than 200 false alarms, the reported rates are lower bounds, and true public tolerance could be even higher than the abstract's headline figures.","The interval-regression design could be combined with a direct elicitation question, such as asking respondents how many false alarms are too many, to test whether respondents hold a stable internal threshold or merely react to the presented trade-off."],"forward_implications":["At the 12.7% air-travel false-positive rate the authors estimate, a fully operational EU-wide PNR system would require on the order of 140,000 manual checks per day, raising the question of whether the accepted error rates are operationally sustainable.","Because respondents found air-travel surveillance more acceptable than sea-travel surveillance even though the privacy consequences are identical, expanding algorithmic screening to new transport modes is likely to face more public resistance than continuing it where it already exists.","The public's acceptance of high false-positive counts suggests that legal challenges to mass surveillance that argue from individual rights may miss the statistical, systemic character of the harm; legitimacy arguments would need to engage with aggregate error rates rather than only individual redress.","Respondents' thresholds did not vary with age or with the number of true positives shown, so within the tested range the acceptability of false positives appears stable across those dimensions."],"supporting_citations":[{"why":"It supplies the contextual-integrity theory that defines the study's measure of legitimacy.","marker":"[46]"},{"why":"It contributes the collective-criticism protocol that turns trade-off answers into interval data for threshold estimation.","marker":"[43]"},{"why":"It quantifies the base-rate consequence: hundreds of thousands of annual false positives even at 99.9% accuracy, motivating the policy question.","marker":"[23]"},{"why":"It sets out the relevant EU court ruling that permits PNR processing with human verification, the legal target of the paper's critique.","marker":"[13]"},{"why":"It argues the PNR ruling ignored the base-rate fallacy, giving the paper its central legal critique.","marker":"[38]"},{"why":"It provides national trust-in-institutions data used to interpret Finland's unusually high acceptance.","marker":"[49]"}],"fun_headline_variants":["Finns' tolerance for travel-screening false positives quantified","Air travel: Finns ok with 13% false alarms, sea lower","Women more tolerant than men of surveillance false alarms","Survey: high false-positive counts seen as legitimate in travel checks","False-positive threshold for travel surveillance: Finns accept ~13%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a person who says security comes first would accept every number of false alarms from the one shown up to 200, and that the statistical curve used to turn those either-or answers into exact thresholds is shaped correctly; if people would accept more than 200, or the curve is misshapen, the reported percentages change.","fun_headline_variants_meta":{"raw":{"variants":["Finns' tolerance for travel-screening false positives quantified","Air travel: Finns ok with 13% false alarms, sea lower","Women more tolerant than men of surveillance false alarms","Survey: high false-positive counts seen as legitimate in travel checks","False-positive threshold for travel surveillance: Finns accept ~13%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1550,"prompt_tokens":994,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":470}},"tokens_in":610,"tokens_out":556,"duration_ms":6004,"temperature":1.0,"reasoning_tokens":470,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:09:04.390887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same vignette with a different denominator, for example 500 or 1,000 flagged passengers, and also ask respondents directly for the maximum acceptable number of false alarms; if the estimate does not scale with the denominator, or direct answers diverge from the regression-based threshold, the reported rates are an artifact of the 200-passenger frame and the assumed curve.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the contextual-integrity theory that defines the study's measure of legitimacy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It contributes the collective-criticism protocol that turns trade-off answers into interval data for threshold estimation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It quantifies the base-rate consequence: hundreds of thousands of annual false positives even at 99.9% accuracy, motivating the policy question."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It sets out the relevant EU court ruling that permits PNR processing with human verification, the legal target of the paper's critique."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It argues the PNR ruling ignored the base-rate fallacy, giving the paper its central legal critique."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides national trust-in-institutions data used to interpret Finland's unusually high acceptance."}],"review_version":1}